Playbook AnswerMonk Research Desk

How to track what AI assistants recommend in your category

TL;DR
  • One prompt is not one answer. In our 300-probe phrasing study, two engine presets agreed on category winners only 22% of the time — a single ChatGPT spot-check is a coin flip, not a measurement.
  • The free protocol. Build 15-20 buyer prompts (5 commercial, 5 category, 5 competitor), run them monthly across ChatGPT, Gemini, and Claude, and record named yes/no, position, and sources cited.
  • Track upstream signals per engine. ChatGPT tracks Bing (~87% citation alignment per Seer Interactive), Perplexity leans on Reddit (~20-24%), and Ahrefs found YouTube mentions the strongest signal at ~0.737.
  • When to use a tool. Spreadsheets break on phrasing variance, engine churn, and source aggregation; a free AnswerMonk audit gives you the cross-engine baseline in 3-8 minutes.

To track what AI assistants recommend in your category, build a fixed panel of 15 to 20 buyer prompts, run it across ChatGPT, Gemini, and Claude once a month, and record three things for every answer: whether your brand was named, where it appeared, and which sources were cited. A spreadsheet handles this fine for a single brand; when you need phrasing variants, engine history, or aggregated source data, that is when a tool earns its keep — and a free AnswerMonk audit gives you the first data point in 3 to 8 minutes.

In July 2026 we asked three engines for tools to track AI recommendations. One answer recommended a tool whose only web presence was a vercel.app preview URL.app preview URL. The assistants do not know how to answer this question yet, which means most businesses are guessing. Below is the protocol we actually use, the free version first.

Why is tracking AI answers different from rank tracking?

Because one prompt does not produce one answer. A rank tracker sends Google the same query every day and gets back roughly the same ranked list, so a single daily check is a fair measurement. An AI assistant regenerates its answer every time it is asked, and small changes in wording, session, or engine configuration change who gets named.

We measured how bad this is. In our 300-probe phrasing study, two different engine presets asked the same questions agreed on category winners only 22% of the time. Same questions, same week, mostly different winners. A single spot-check — you type one prompt into ChatGPT, see a competitor, and panic — is not a measurement. It is one draw from a distribution.

Wording changes more than the winner; it changes whether the engine looks anything up at all. Nearly a third of the 2,994 answers in our July 2026 study (30.1%) carried no sources at all — pure model memory. And the trigger can be a single word. Identical vet-clinic questions about Toronto that omitted the city name got memory-only answers 81% of the time; add "Toronto" to the question and that dropped to 0% — every answer retrieved and cited live sources. Austin gyms showed the same flip, 75% versus 0%.

So the unit of measurement is not a prompt. It is a panel of prompts, run on a schedule, with results recorded whether you like them or not. That is the core mental shift from SEO, and if the broader difference between the two disciplines is new to you, start with our GEO vs SEO explainer and come back.

How do you build a free manual tracking panel?

Write 15 to 20 prompts a real buyer would type, split three ways: five commercial, five category, five competitor. Do not write them the way a marketer talks. Write them the way a stressed person with a problem talks.

  • Commercial prompts (5): "best [service] in [city]", "top [product] for [specific use case]", "[service] near me that does [thing]". Include your city or region by name — per the city-name effect above, a named location almost always flips the engine into live retrieval in our data (Toronto vet questions went from 81% memory-only to 0%; London plumbers from 55% to 6%), which is the answer mode where local businesses actually get cited.
  • Category prompts (5): "how do I choose a [service]", "is [treatment/product] worth it", "what does [service] cost". These often get memory-only answers, and that is the point: they tell you what the model already believes about your category before it looks anything up.
  • Competitor prompts (5): "[Competitor] alternatives", "[Competitor] vs [you]", "[Competitor] reviews". You are testing whether engines volunteer your name when a buyer starts from someone else's.

Run the panel monthly. Semrush's September 2025 longitudinal test tracked 81 new pages for 30 days: 42% were cited by ChatGPT within 30 days, and the rest mostly never were. AI answers appear to move on a scale of weeks rather than hours — that is one month of Semrush data, so treat it as directional, but it is the best cadence evidence available. Monthly is frequent enough to catch real movement and infrequent enough that you will actually keep doing it.

Record results in a spreadsheet with one row per prompt and one column group per engine. Each group needs exactly three columns:

  1. Named: yes or no. Was your brand mentioned anywhere in the answer?
  2. Position: first named, mentioned later, or absent. Being the first recommendation is a different outcome from being fourth in a list.
  3. Sources: the domains cited, comma-separated, or "none" for a memory-only answer.

The sources column is the one people skip, and it is the most valuable. In our 200-probe Dubai dental set, engines averaged 11.8 citations per answer, and the most-cited sources were clinics' own websites — not health portals. Whatever domains keep showing up in your category's citations are the places your next quarter of content and outreach effort should go.

Two hygiene rules. First, use a fresh chat for every prompt, logged out or with memory and personalization disabled where the engine allows it — a personalized session shows you an answer no new buyer will ever see. Second, run each prompt once per engine and record what comes back. Rerunning until you get an answer you like is the fastest way to make the whole exercise worthless.

What should you track per engine?

Track the upstream signal each engine leans on, not just the answer text. The four major assistants retrieve from different substrates, so the same visibility gap has a different cause — and a different fix — depending on which engine you are losing in.

ChatGPT: watch Bing. ChatGPT Search is, in practice, a Bing surface: Seer Interactive puts the citation overlap with Bing's top organic results at about 87%. If ChatGPT will not name you, check your Bing rankings for your panel's commercial prompts before you change anything on your site. Note that domain authority is not the gate here: industry analyses find around 90% of pages ChatGPT cites rank on Google page 3 or beyond. Our ChatGPT traffic playbook covers the fix side in detail.

Perplexity: watch Reddit. Industry citation analyses put Reddit at roughly 20-24% of Perplexity's citations. Add a fourth tracking column for Perplexity: whether your brand appears in the subreddit threads its answers cite. If your competitors live in those threads and you do not, that is the gap.

Gemini and Google's AI surfaces: watch mentions, especially YouTube. In Ahrefs' 75,000-brand dataset, mentions beat backlinks three-to-one as predictors of AI visibility (0.664 against 0.218), and YouTube mentions top the list at roughly 0.737. Track where your brand is mentioned each month, not just where it is linked. Muck Rack's citation research points the same direction: 84% of AI citations come from earned media.

Software and services categories: watch G2 and Capterra. In our July 2026 live probes, G2 and Capterra were recommended in nearly every answer about getting mentioned by ChatGPT. If you sell software, your review-site profiles are part of your AI footprint whether you maintain them or not, so put them in the spreadsheet too.

When does a spreadsheet stop scaling?

The manual panel breaks on three things: phrasing variance, engine churn, and source aggregation. None of them break it immediately. For a single local brand, the protocol above genuinely works, and we would rather you run it than wait for budget.

Phrasing variance is the first wall. Twenty prompts is one sample per question, and our phrasing study showed how differently the same question can land. Smoothing that noise properly means three to five wording variants per prompt — 60 to 100 runs per engine, per month, by hand. That is a full working day of copy-paste.

Engine churn is the second. Models and presets update without notice. When your appearance rate drops in October, you cannot tell whether your visibility fell or the engine changed under you — unless you have enough run history to separate the two, which a monthly 20-prompt sheet does not give you.

Source aggregation is the third. Reading citations one answer at a time works; answering "which ten domains drive recommendations across my whole category, on every major engine, and am I on any of them" does not fit in a spreadsheet at all.

ApproachWhat you getCostWhere it breaks
Manual prompt panelNamed yes/no, position, and basic sources for 15-20 promptsFree; 2-3 hours a monthPhrasing variance, no history at volume, no source aggregation
General SEO toolsGoogle (and sometimes Bing) rank trackingYour existing subscriptionThey measure the ranked-list machine; AI-answer coverage is early and partial
Dedicated AI-visibility toolsAppearance rates across engines, competitor leaderboards, citation-source breakdownsFree audits to paid monthly plansYoung category — vet vendors before you buy

That last cell is not a throwaway. Remember the probe that opened this piece: an engine recommending a tracking tool that existed only as a vercel.app preview URL. This category is young enough that vendor diligence matters, which is why we published a full ranked comparison in our best AI visibility tools guide. Tools buy you scale, run history, and source data. They do not buy you magic, and any vendor implying otherwise has told you something useful about themselves.

What is the fastest way to get a first data point?

Run the free AnswerMonk audit: enter your URL at answermonk.ai, no signup, results in 3 to 8 minutes. It generates the buyer prompts in your category, runs them across ChatGPT, Gemini, and Claude, and returns an appearance rate, a competitor leaderboard, and a citation-source breakdown — the same three columns as the spreadsheet, plus the aggregation the spreadsheet cannot do.

The sensible workflow is both. Use the audit as your day-zero baseline, then keep the manual panel running monthly against it; the audit tells you where you stand across engines, and the panel keeps you close to the actual answer text buyers see. If you want the measurement turned into a to-do list, the $19/month tier adds plain-language action plans, a living knowledge base tracking Schema.org, Google, OpenAI, and Perplexity guidance, a source-placement agent, and WhatsApp lead capture. You can see what movement looks like in practice in the public before/after reports at /reports.

One honest closing note. Our own numbers here come from July 2026 probe data, and the external studies we cite are each single snapshots too — one snapshot of fast-moving systems, so treat the exact percentages as directional. The finding we are confident in is the structural one: nobody owns the answer to "who does AI recommend in my category," not even the engines themselves. One month of measurement will not change your answers, but a year of it will tell you what actually moved them — and almost nobody in your category has that record yet.

Run your free AI visibility audit →


Frequently asked questions

How often should I check what ChatGPT recommends in my category?

Monthly is the right cadence for most brands. Semrush's September 2025 longitudinal test found 42% of new pages were cited by ChatGPT within 30 days, and the rest mostly never were, so answers move on a scale of weeks rather than days. Checking more often mostly adds noise from answer-to-answer variance.

Can I track AI answers with a normal rank tracker like Semrush or Ahrefs?

Not directly. Rank trackers measure positions in a ranked list of links, while AI assistants synthesize a single answer and usually name only one or two brands. Some SEO suites are adding AI-answer features, but the core product measures a different machine; you need to capture the answer text and its citations, not a position.

Why does Perplexity recommend different businesses than ChatGPT?

They retrieve from different substrates. Industry citation analyses put Reddit at roughly 20-24% of Perplexity's citations, while Seer Interactive found about 87% of ChatGPT Search citations align with Bing's top organic results. A brand that is strong on Bing but absent from Reddit can win one engine and lose the other.

Should I log out of ChatGPT when running tracking prompts?

Yes, or at minimum use a fresh chat with memory and personalization turned off. ChatGPT adapts answers to your history, so a logged-in session can show you an answer no real buyer would ever see. Run each prompt once in a clean session and record whatever comes back, even if you do not like it.

How many prompts do I need to measure visibility in ChatGPT and Perplexity reliably?

Treat 15 to 20 as the floor, not the target. In our 300-probe phrasing study, two engine presets asked the same questions agreed on category winners only 22% of the time, so any small panel is a directional sample rather than a scoreboard. More prompts and phrasing variants tighten the estimate; a single spot-check tells you almost nothing.

Is AnswerMonk's AI visibility audit actually free?

Yes. You enter your URL at answermonk.ai, there is no signup, and results arrive in 3 to 8 minutes. The audit generates the buyer prompts in your category, runs them across ChatGPT, Gemini, and Claude, and returns your appearance rate, a competitor leaderboard, and a citation-source breakdown.

See where your brand ranks in AI search
Enter your domain and get a free AI visibility audit across ChatGPT, Gemini, and Claude.
Run a free audit →