Research AnswerMonk Research Desk

We Analyzed 893 AI Citations: What Gets Cited (2026)

TL;DR
  • Recency is a content format. Of the 547 citations in our July 2026 capture that carried a page title, 36.2% were year-stamped 2025/2026 — the strongest title pattern we measured, ahead of "best" listicles (16.6%) and how-tos (16.1%).
  • Nobody owns AI search yet. The 806 citations with a resolvable domain spread across 490 unique domains, and 357 of those domains were cited exactly once — one-off domains that between them took 44% of every citation we could resolve. Mean top-3 domain share per prompt: 22.2%.
  • The three engines cite like three different people. ChatGPT leaned on official documentation — help.openai.com 21 times — plus bing.com (8) and Reddit (7); Claude cited niche vendor blogs freely, including trysight.ai 10 times from Claude alone; Claude and ChatGPT together cited arxiv.org 18 times.
  • Vs-pages are a graveyard. "X vs Y" titles earned 0.9% of titled citations. "Alternatives"-titled pages: 1.5%. Numbered lists earned 11.9% and "best" listicles 16.6%. When buyers ask comparison questions, engines cite lists.
  • Our own score was 0 of 893. answermonk.ai appeared nowhere in this dataset. This post is us following our own data, in public.
  • See your category's citation map. Start with the free 3–8 minute AI visibility audit — no signup; the paid tier is $19/mo, and the same fan-out capture, cited-URL analysis, and authority-domain identification behind this study run for your category, with public before/after reports.
  • In July 2026 we captured the fan-out queries and citations behind 24 real buyer prompts across ChatGPT, Claude, and Gemini — 893 citations in total. Four findings held across the dataset: recency behaves like a format of its own (36.2% of titled citations were year-stamped), fragmentation beats authority (490 unique domains, 73% of them cited exactly once, mean top-3 share per prompt just 22.2%), the three engines have distinct citation personalities, and "X vs Y" comparison pages are nearly uncitable at 0.9%. One more finding, about us: our own domain took 0 of the 893. Every number below comes from that single dated capture — and you can see the same map for your own category with a free 3–8 minute AI visibility audit, no signup.

What content format gets cited most by AI search engines?

Year-stamped content. Of the 547 citations in our capture that carried a page title, 198 — 36.2% — had 2025 or 2026 in the title. That is more than double the share of any classic content format. Recency is not a tiebreaker in AI search; it behaves like a format in its own right — and not only in our data: Semrush's answer-engine research ties visible timestamps and recency to citation frequency.

Here is the full pattern breakdown. Percentages are of the 547 titled citations — the remaining rows carried no page title:

Title patternCitationsShare of 547 titled citations
Year-stamped (2025/2026)19836.2%
"Best" listicle9116.6%
How-to8816.1%
Guide / playbook / checklist7213.2%
Numbered list6511.9%
Question-form title274.9%
"Alternatives" in title81.5%
"X vs Y"50.9%

Patterns overlap — a single title can be both year-stamped and a numbered list — so the shares are not meant to sum to 100%.

The recency finding staged its own demonstration inside the dataset. When we ran the prompt "What content format gets cited most by AI search engines?", the engines fanned out into retrieval queries that included, verbatim, "what content type do AI search engines cite most 2026 study". The retrieval layer was searching for a year-stamped study. Nothing study-shaped won the slot — of the 42 citations that came back for this prompt, the single most-cited source was reddit.com, with 4, threads filling a vacuum where research should be. The page you are reading is our attempt to be that study, year stamp included.

One reconciliation, because the 4.9% question-form number looks like it contradicts advice we give elsewhere: question-form page titles are rare among citations, but question-form structure inside pages is a separate lever. Our earlier crawl of citation-winning pages found FAQ blocks 3.4x more likely to be cited and question-form headings roughly 2x. Titles sell recency and list-shape to the retrieval layer; body structure sells extractability once the page is fetched. GEO vs SEO covers that split in depth.

Which domains do AI engines actually cite — and how fragmented is it?

806 of the 893 citations resolved to a domain, and those 806 spread across 490 unique domains — 357 of them, or 73%, cited exactly once. Citations to those one-off domains add up to 44% of everything we could resolve. The mean top-3 domain concentration per prompt was 22.2%. There is no Healthline of this category — no single publisher the engines reflexively reach for. Nobody owns it.

The range around that mean runs from 9% to 52%, and the extremes are instructive. The most concentrated prompt in the set was the definitional one — "What is generative engine optimization and does it work?" — at a 52% top-3 share, led by arxiv.org and developers.google.com. Definitional questions have canonical sources: the academic paper that coined the term, the platform documentation that governs it. Commercial questions scatter. "Which AI visibility tools show evidence for their recommendations?" was the most fragmented prompt in the set at an 8.8% top-3 share, and "Do backlinks matter for AI search visibility?" — a question the entire SEO industry argues about — topped out with its three most-cited domains at two citations each, a 20% top-3 share with no recognized authority behind it.

An honesty note on that backlinks prompt: our capture shows who gets cited when the question is asked, not whether backlinks cause citations. That's a correlation study we have not run, and this dataset can't answer it.

Fragmentation cuts both ways. A long tail where one-off domains soak up 44% of resolvable citations means the citation slots in most commercial prompts are genuinely open — small domains took them throughout our capture. It also means today's winners hold shallow moats. If a competitor of yours is being cited right now, treat it as a lead to reverse-engineer rather than a verdict — we wrote up how to read a competitor's citations. And the openness isn't unique to our capture: Ahrefs measured that ~62% of AI Overview citations come from pages outside the top-10 organic results.

Do the three engines cite differently?

Yes — differently enough that measuring one engine and extrapolating is a sampling error. Our capture returned 390 citation rows from Claude, 346 from Gemini, and 157 from ChatGPT, and the sources behind those rows barely resemble each other.

EngineCitation rowsCitation habits in our capture
Claude390Cites niche vendor blogs freely — trysight.ai drew 10 citations from Claude alone; shares heavy arxiv.org use with ChatGPT, 18 citations between the two engines
Gemini346Second to Claude in row count, but 87 of those rows are query-only — Gemini's API doesn't map citations to individual fan-out queries
ChatGPT157Official documentation first — help.openai.com 21 times; bing.com 8; reddit.com 7

Two cross-engine notes: Claude and ChatGPT together cited arxiv.org 18 times, and developers.google.com took 12 citations across the dataset — platform documentation and academic sources travel across engines in a way most marketing content does not.

The practical consequence: the page that wins Claude citations — deep, specific vendor content in a niche — is not automatically the page that wins ChatGPT, which in our capture favored official docs, indexed pages, and forums. A brand can look visible on one engine and be absent on another while both answer the same buyer question. Measure per engine, or you're guessing.

What never gets cited — and why is the vs-page graveyard so full?

"X vs Y" comparison pages earned 5 citations out of 547 titled rows — 0.9%, the lowest share of any pattern we classified. "Alternatives"-titled pages did barely better at 1.5%. Meanwhile numbered lists took 11.9% and "best" listicles 16.6%. When a buyer asks a comparison-shaped question, the engines overwhelmingly cite pages structured as lists of options — not head-to-head duels.

That should reorder a lot of content roadmaps. The "YourBrand vs Competitor" page is a staple of SEO programs, and in our capture it is nearly invisible at the retrieval layer. The same comparison material, restructured as a year-stamped list with explicit criteria, matches the shape the engines actually cited. Write the listicle.

And then there is the deepest layer of the graveyard: domains that never appear at all. Ours, for one. answermonk.ai was cited 0 times out of 893. We build AI visibility software, and in July 2026 we were invisible at the fan-out layer of our own category. We're publishing that number instead of burying it because it is the entire reason this research program exists: the dataset told us what gets cited — year-stamped, list-shaped, question-structured research — so we are shipping exactly that, and we'll publish the re-audit delta at /reports whether it flatters us or not.

Methodology — how did we capture the fan-out layer?

We recorded, for each of 24 real buyer prompts, every fan-out query each engine issued at the retrieval layer, plus every URL the engine cited, with rank and attribution level. Here is why that layer is the one worth capturing: when you ask an AI engine a buyer question, it doesn't search your words. It fans out into multiple retrieval queries, fetches sources for each, and assembles its answer from what came back. That fan-out layer — not the visible answer — is where citation decisions happen.

The 24 prompts spanned tool questions like "Best tools to track whether AI assistants recommend my business", diagnostic ones like "Why is my competitor cited by ChatGPT and not me?", and vertical ones from dental clinics to hotels.

The dataset:

  • 24 prompts, 3 engines, captured July 2026
  • 893 citation rows: Claude 390, Gemini 346, ChatGPT 157
  • 806 rows resolved to a domain; 547 carried a page title
  • Attribution: 536 rows tie to a specific fan-out query, 270 tie to the prompt level, and 87 are query-only

The limitations, plainly:

  1. One category cluster. All 24 prompts orbit AI search visibility and the businesses asking about it. Citation behavior in your category may differ — that's an argument for measuring it, not for extrapolating from ours.
  2. One run. This is a July 2026 snapshot. Engines change retrieval behavior without notice, and we make no claim these numbers hold in December.
  3. Gemini attribution caveat. Gemini's API doesn't map citations to individual fan-out queries, so 87 Gemini rows are query-only, and some Gemini grounding redirects yielded no resolvable domain.
  4. Denominators matter. Title-pattern percentages use the 547 titled rows as denominator, not all 893. Where we say 36.2%, we mean 36.2% of titled citations. Per-prompt top-3 shares are computed over the rows that resolved to a domain for that prompt, not all captured rows — recomputing against raw cite counts will read slightly low.
  5. First-party instrument. Disclosure: AnswerMonk is our product, and this dataset was captured with AnswerMonk's own fan-out capture layer — the same capture it runs per category. The numbers above carry their denominators and their capture date so you can judge for yourself; our bar for what "showing evidence" should mean in this category is written up in AI visibility tools that show their evidence.

How do you run this analysis for your own category?

Start free, and only pay if the map is worth acting on. The free AI visibility audit takes 3–8 minutes, no signup: it returns your appearance rate, a competitor leaderboard, and a citation-source breakdown for your category. The paid tier is $19/month. The toolkit behind this study — fan-out prompt capture, what-works-in-your-category analysis, cited-URL analysis, authority-domain identification, and re-audit delta verification after you ship fixes — is the same analysis AnswerMonk runs for your category. What it does not include is a guarantee: nobody can honestly promise you a spot in AI answers, and we don't. Public before/after reports live at /reports — and our own re-audit delta from this 0-of-893 baseline will be published there, whichever way it moves.

Run your free AI visibility audit →


Frequently asked questions

What content format gets cited most by AI search engines?

Year-stamped content, in our July 2026 capture of 893 citations across ChatGPT, Claude, and Gemini. Of the 547 citations carrying a page title, 36.2% had 2025 or 2026 in the title — ahead of "best" listicles (16.6%), how-tos (16.1%), guides (13.2%), and numbered lists (11.9%). Recency behaves like a format, not a bonus signal.

How fragmented are AI search citations?

Very. The 806 citations with a resolvable domain spread across 490 unique domains, and 357 of them were cited exactly once; 44% of resolvable citations went to a domain that appeared only that once. The mean top-3 domain share per prompt was 22.2%, ranging from 9% on the most scattered commercial prompt to 52% on the definitional "what is GEO" prompt. In most commercial prompts, no source has locked up the answer.

Do ChatGPT, Claude, and Gemini cite the same sources?

No. In our capture, ChatGPT leaned on official documentation (help.openai.com 21 times), bing.com, and Reddit; Claude cited niche vendor blogs freely — one vendor blog drew 10 citations from Claude alone — and shares heavy arxiv.org use with ChatGPT, 18 citations between the two engines; Gemini produced the caveat that its API doesn't map citations to individual fan-out queries. A brand can be visible on one engine and absent on another for the same question.

Are "brand vs brand" comparison pages worth writing for AI search?

Our data says no — not in that shape. "X vs Y" titles earned 0.9% of titled citations, the lowest of any pattern we classified, while "best" listicles earned 16.6% and numbered lists 11.9%. The same comparison material restructured as a year-stamped list with explicit criteria matches what engines actually cited.

Why publish a study that shows your own tool at zero citations?

Because it's true, and because it's the point. answermonk.ai took 0 of the 893 citations in our own capture. That zero is why this research program exists: the dataset shows what gets cited, we're following it in public, and the re-audit delta will be published at /reports either way. We'd rather show the before than pretend there wasn't one.

Can I see this citation data for my own category?

Yes. The free AnswerMonk audit — 3 to 8 minutes, no signup — returns your appearance rate, competitor leaderboard, and citation-source breakdown. The paid tier is $19/month, and the same fan-out capture behind this study runs for your category: the retrieval queries engines run, the URLs they cite, the domains that carry authority there, and re-audit verification after fixes.

See where your brand ranks in AI search
Enter your domain and get a free AI visibility audit across ChatGPT, Gemini, and Claude.
Run a free audit →