Skip to content

4 min read

Stage 01 · Demand: Being Findable When the Searcher Is a Machine

How retrieval actually selects sources, what the first peer-reviewed study of generative engine optimization found, and why AI search cites so few domains.

  • SEO
  • AI Visibility
  • Retrieval

A growing share of your buyers will never see a list of ten blue links. They will read one synthesised answer with a handful of citations. The question stops being "do I rank" and becomes "am I in the retrieved set, and am I quotable."

Ariadne's first stage is Demand, and the site describes it as "your expertise, clearly cited." That phrasing is deliberate, and this post is the argument behind it.

How a machine decides you exist

Two mechanisms, working together.

The first is lexical. BM25 remains the default sparse ranking function and the sparse half of essentially every hybrid retriever in production. Robertson and Zaragoza's monograph is the definitive derivation: term frequency with saturation, inverse document frequency, and document-length normalisation, governed by free parameters k1k_1 and bb Robertson & Zaragoza 2009. The scoring form is

BM25(q,d)  =  tqIDF(t)f(t,d)(k1+1)f(t,d)+k1(1b+bdd)\mathrm{BM25}(q,d) \;=\; \sum_{t \in q} \mathrm{IDF}(t)\cdot \frac{f(t,d)\,(k_1+1)}{f(t,d) + k_1\left(1 - b + b\,\frac{|d|}{\overline{|d|}}\right)}

The saturation term matters practically: repeating a keyword a tenth time adds almost nothing, because f(t,d)f(t,d) enters through a function that flattens. That is the mathematical reason keyword stuffing stopped working, stated precisely rather than as folklore.

The second mechanism is dense. Karpukhin and colleagues showed a simple dual-encoder trained on a modest number of question–passage pairs beats BM25 decisively — by 9 to 19 percentage points absolute in top-20 passage retrieval accuracy, with a headline single-dataset result of 78.4% against 59.1% Karpukhin et al. 2020. This is why a page that never uses the searcher's exact words can still be retrieved: its embedding is close in meaning.

Retrieval-augmented generation stitches these together. Lewis and colleagues introduced the architecture that now underlies AI answers: a generator conditioned on passages pulled from a dense index by a neural retriever, producing "more specific, diverse and factual language" than a parametric-only baseline Lewis et al. 2020.

Figure 1The path from a buyer's question to a cited answer. You are competing for a slot in the retrieved set, and then again for a citation inside the generated answer.Diagram by the author.

What actually moves visibility in generative answers

The first serious academic treatment of this question is the GEO paper, published at KDD 2024 Aggarwal et al. 2024. The authors built GEO-bench — 10,000 queries drawn from nine datasets across 25 domains — and tested which content modifications increase the chance of being cited in a generative engine's response.

The results are unusually actionable, and unusually inconvenient for traditional SEO practice:

  • Adding citations, direct quotations and statistics produced the largest gains, in the range of 30–40% relative visibility improvement.
  • Keyword stuffing underperformed — the classic lever transfers poorly.
  • The effect is levelling rather than rich-get-richer: lower-ranked sites gained proportionally more.

The uncomfortable part: concentration and error

Two findings should temper any promise anyone makes you about "AI visibility."

Citations concentrate hard. Kai-Cheng Yang measured citation behaviour across OpenAI, Perplexity and Google models using AI Search Arena data — over 24,000 conversations, 65,000 responses and 366,000 citations. Among news citations, the top 20 sources accounted for 67.3% of the total Yang 2025. The realistic goal is entering a small cited set within a narrow topic, not general visibility.

The engines cite badly. Liu, Zhang and Liang performed the first rigorous human audit of generative search citation quality. Averaged across four engines, only 51.5% of generated sentences were fully supported by their citations, and only 74.5% of citations supported the sentence they were attached to Liu et al. 2023. Their most counterintuitive result: citation precision is strongly inversely correlated with perceived utility, at r = −0.96. The answers users like best are the ones whose citations support them least.

The Tow Center's audit found related pathologies at the publisher level: across 1,600 queries, the tools gave incorrect answers to more than 60%, fabricated links, and frequently cited syndicated copies rather than the original publisher Jaźwińska & Chandrasekar 2025.

The structured-data layer

One more substrate worth getting right. Schema.org was founded in 2011 by Bing, Google and Yahoo (Yandex joined later) as a shared vocabulary for machine-readable page meaning; by the time its creators wrote it up in CACM, markup appeared on 31.3% of pages measured across roughly 10 billion pages, up from 22% a year earlier Guha et al. 2016.

Structured data is not a ranking trick. It is the difference between a machine inferring what your page is about and being told. For a service business, the entities worth declaring are unglamorous and specific: organisation, service, area served, review, FAQ, and the person who does the work.

The stage in one sentence

Demand work is the discipline of making your expertise machine-legible and quotable — measured not by rankings, but by whether the number of qualified conversations went up.

References

  1. Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019
  2. Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense passage retrieval for open-domain question answering. Proceedings of EMNLP 2020, 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550
  3. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). https://arxiv.org/abs/2005.11401
  4. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. Proceedings of KDD 2024, 5–16. https://doi.org/10.1145/3637528.3671900
  5. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. Findings of ACL: EMNLP 2023, 7001–7025. https://aclanthology.org/2023.findings-emnlp.467/
  6. Yang, K. (2025). News source citing patterns in AI search systems. arXiv preprint arXiv:2507.05301. https://arxiv.org/abs/2507.05301Preprint; not peer-reviewed at time of writing.
  7. Jaźwińska, K., & Chandrasekar, A. (2025). AI search has a citation problem. Columbia Journalism Review / Tow Center for Digital Journalism. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.phpJournalism-institute research, not peer-reviewed academic work.
  8. Guha, R. V., Brickley, D., & Macbeth, S. (2016). Schema.org: Evolution of structured data on the web. Communications of the ACM, 59(2), 44–51. https://doi.org/10.1145/2844544

Stage 02 is next: what happens once the machine has sent you a human.

Sourena Khanzadeh

Founder & Growth Engineer, Ariadne Growth Systems

Toronto, Canada

Ariadne Growth SystemsGrowth System Auditsupport@ariadne.fyi