How AnswerMonk measures AI visibility

AnswerMonk measures AI visibility by generating the buyer questions real customers ask in your category, running them live against ChatGPT in three search modes, recording every answer and citation, and reporting the results as evidence-backed ranges rather than a single score. This page explains what an audit actually runs, why we report the way we do, and where the limits are. The audit itself is free, requires no signup, and takes 3-8 minutes.

What an audit actually runs

An audit starts with prompt generation. For your category, we build a panel of buyer questions in three families: commercial questions ("best CRM for a 10-person sales team"), category questions ("how do I choose a CRM?"), and competitor questions ("X vs Y — which should I pick?"). These are the phrasings that put brands into answers, not vanity prompts about your own name.

Each prompt runs live against ChatGPT in three search modes: search off, model decides, and search forced. We record the full answer each run gives, including whether it retrieved from the web or answered from memory, and we extract every citation it attaches. Nothing is simulated or backfilled from historical data — every data point in your report is a real answer from a real engine, captured at run time.

From those recorded runs the report assembles your appearance rate, a competitor leaderboard, a citation-source breakdown, per-mode leader comparisons, a winners playbook drawn from what cited brands do differently, and a claims checker that flags what ChatGPT gets wrong about you.

Why we report ranges, not a single score

AI answers are not stable enough for one number to be honest. In our 300-probe phrasing study, two engine presets agreed on category winners only 22% of the time — the same question, phrased or configured differently, crowned different brands. And in our 2,994-probe July 2026 study across 19 categories, 30.1% of answers cited zero sources: the engine answered purely from memory, where training data rather than today's web decides who gets named.

Small wording changes flip behavior too. When we asked identical local-service questions with and without a city name, 81% of answers were memory-only without the city and 0% with it (Toronto vets; the same flip held for Austin gyms, London plumbers, and Sydney accountants). A single spot-check — asking ChatGPT one question, once — can land anywhere in that spread and tell you almost nothing. So we run panels of prompts, report appearance as a rate across the panel, and show per-mode differences instead of collapsing them into one flattering average.

The evidence chain

Every recommendation in an AnswerMonk report traces back to something you can open: the recorded prompt, the answer ChatGPT gave, and the source it cited. If the playbook says review platforms drive citations in your category, the report shows the answers where that happened. If the claims checker flags that ChatGPT states your pricing wrong, the flag carries per-claim provenance — the exact answer, the claim within it, and the source (or absence of one) behind it.

We built it this way because unverifiable advice is the norm in this space, and we did not want to add to it. You should never have to take a chart on faith.

Our research corpus

Recommendations are only as good as the data behind them, so we run our own studies continuously. The core corpus is a 2,994-probe study across 19 categories, re-run weekly, tracking how often engines cite sources, which domains win citations, and how category leaders shift over time. Findings from the corpus — such as which page structures over-index among cited sources — feed directly into the playbooks in customer reports. We publish public examples and standing insights at /reports, so you can inspect the kind of output an audit produces before running one.

Our studies, named and dated

Where an article of ours cites a first-party figure, it comes from one of the studies below. Each is dated, with its method and sample size, so you can judge the claim instead of taking it on trust.

Sample size decides how far a number can be trusted. A rate measured over 30 runs carries roughly a ±17-point confidence interval; over 50 runs, about ±13. We are moving published figures to state the sample and the interval next to the estimate, and to report per-engine results rather than one blended score.

Limitations we publish

Three limits are worth stating plainly. First, one-month windows are directional, not definitive: engines update on their own schedules, and a shift in your numbers can reflect their changes as much as yours. Second, engines change — models get swapped, retrieval behavior gets tuned, and past patterns are not guarantees of future ones. Third, our probes are samples. A panel of buyer questions covers the phrasings we can generate and test, not the full universe of ways a real buyer might ask. We design panels to be representative, but representative is not exhaustive, and we would rather say so than imply otherwise.

If a finding survives those caveats — it holds across engines, phrasings, and reruns — it goes in your report. If it does not, it stays in the lab. Run the free audit at answermonk.ai to see the method applied to your own category.

Frequently asked questions

Which AI engine does AnswerMonk test?

Every audit runs your prompt panel on ChatGPT across three search modes — search off, on-demand, and forced. Answers and citations are recorded per mode, so the report shows where each one differs instead of blending them into a single average.

How long does an AnswerMonk audit take?

Between 3 and 8 minutes, depending on category size. It is free and requires no signup. You get an appearance rate, a competitor leaderboard, a citation-source breakdown, per-mode comparisons, a winners playbook, and a claims checker.

Why doesn't AnswerMonk report a single AI visibility score?

Because the underlying data is too variable for one number to be honest. In our 300-probe phrasing study, two engine presets agreed on category winners only 22% of the time. We report rates across a prompt panel and per-mode differences instead.

Can I see an AnswerMonk report before running my own audit?

Yes. Public example reports and standing research insights are published at answermonk.ai/reports. They show the same structure and evidence chain a report on your own category would contain.

Methodology