Key findings
- Engines read the site more often than they name it. A buzzboxmedia.com URL appeared in the engine's own source list on 117 of 195 calls, which is 60.0%. The brand appeared in the answer prose on 88, which is 45.1%. The gap runs one way only: 29 calls read a page without naming us, and zero calls named us without reading a page.
- Half of every apparent win is unstable. Each question was put to each engine five times with identical wording on the same day. Of the 22 question and engine pairs that named the brand at least once, 11 did not do it on every run. Citation as a source is steadier but still moves: 7 of the 26 pairs that ever cited a page failed to do so on all five runs.
- Citation concentrates into almost nothing. 22 pages out of 1,158 produced all 156 page citation events. One roundup article took 41 of them, or 26.3%. The top three pages took 41.7%. 290 conference pages, a quarter of the site, were cited zero times.
- Question type matters far more than any page attribute we can measure. On the five agency-recommendation questions, engines cited a page on 80.0% of calls and named us on 70.7%. On the eight informational questions, 47.5% and 29.2%.
- Formatting can be retrofitted across a whole site in a week. Evidence cannot. Between two commits six days apart, scored by the same program, question-shaped headings went from 76.9% to 94.8% of pages, FAQ schema parity from 74.2% to 90.8%, and a visible freshness signal from 10.4% to 91.8%. The two criteria that require sourcing rather than structure did not move: extractable facts with sources went from 37.0% to 36.6% and claims integrity from 32.6% to 33.1%.
- We could not reproduce our own earlier finding. An internal audit of ours in early August reported that AI-cited pages fail claims integrity at roughly twice the site rate. The direction repeats here, 18.2% against 33.4%, and at 22 cited pages it does not reach significance, Fisher exact two-sided p = 0.17. We are publishing it as an open question rather than as a finding, and correcting the earlier version.
Being read and being recommended are different outcomes
Most AI visibility checks ask one question: does the brand name appear in the answer. That test is easier to get wrong than it looks, and getting it wrong is what an earlier version of this study did.
Answer engines write their citations inline, as links. So the moment an engine cites your page, your own domain lands in the answer text, and a naive search for your brand string inside that text starts returning true for reasons that have nothing to do with being recommended. On our first 71 calls, the naive test returned identical results to "was a URL cited" on 71 of 71, and 11 of its 45 positives had no brand mention anywhere outside a URL.
So this study measures two outcomes separately and never merges them:
- Cited as a source. A buzzboxmedia.com URL is in the structured source list the engine's API returns. The engine read the page.
- Named in the prose. The brand appears in the answer text after every URL, markdown link target and bare host token has been stripped out. The engine told the reader our name.
| Engine | Calls | Cited a page | Named the brand | Sources per call |
|---|---|---|---|---|
| ChatGPT | 65 | 67.7% | 43.1% | 5.0 |
| Claude | 65 | 55.4% | 41.5% | 8.7 |
| Grok | 65 | 56.9% | 50.8% | 4.8 |
| All three pooled | 195 | 60.0% | 45.1% | 6.2 |
Pooled, the 95% confidence interval on the citation rate is 53.0% to 66.6% and on the naming rate 38.3% to 52.1%. Those intervals overlap the individual engines' intervals in every case, so we are not claiming that any one engine cites us more than another. What we are claiming is the direction of the gap within each engine, which holds for all three.
The practical consequence is that a page can be doing its job and be invisible in the check most people run. Twenty-nine times, an engine went to one of our pages, took what it needed, and answered without attribution the reader could act on. If you are only counting brand mentions, that work registers as zero.
The same question, asked five times
This is the part we did not expect to be the headline.
Every published AI visibility audit we have seen, including the one we sell, asks each question once. We asked each question five times, per engine, on the same day, with identical wording, and treated the five as independent observations.
| Outcome | Same on all 5 runs | Changed between runs | Never, all 5 runs |
|---|---|---|---|
| Named in the answer | 11 | 11 | 17 |
| Cited as a source | 19 | 7 | 13 |
Read the naming row the way a marketer would have to. Twenty-two pairs produced a brand mention at some point across the five runs. Eleven of those produced it every time. The other eleven would have reported as a win or as a loss depending purely on which run you happened to catch.
The instability is not spread evenly. ChatGPT changed its mind on 1 of its 13 questions, Claude on 5, and Grok on 5. We are not going to build a theory on 13 questions per engine, but the difference is large enough that anyone tracking a single engine and generalising should be careful.
Two things follow, and both are uncomfortable for how this category currently sells itself.
A one-run audit is not a measurement. It is one draw from a distribution nobody has characterised. That includes before-and-after tests: if you check a query, change the page, and check again, you have two single draws, and on these numbers there is a meaningful chance the difference you observe is the same page behaving differently.
Five runs is not enough either. We want to be honest about our own instrument. With five runs, a pair whose true per-run probability is exactly one half still looks perfectly stable 6.25% of the time, and a pair sitting at 80% looks like a clean sweep about a third of the time. So our 11 stable pairs almost certainly contain some pairs that are not really stable. The direction of that error makes the instability we report an undercount, not an overcount.
Which pages actually get cited
Across 195 calls the engines produced 156 page citation events, spread across 22 distinct pages. The site has 1,158.
The distribution has a very short head. One page, a roundup of medical device marketing agencies, accounts for 41 of the 156 events, or 26.3%. The top three account for 41.7%. Thirteen of the 22 cited pages were cited five times or fewer across the entire run.
Page class matters more than page count. Sixteen guides produced 5 cited pages. Seven hundred and eighty-nine blog posts produced 12. Two hundred and ninety conference pages, a quarter of the entire site, produced none. That last number is worth sitting with, because the conference pages are the ones our formatting retrofit improved most.
The engines also disagree about the head. Of the 22 cited pages, 14 were cited by only one of the three engines, 5 by two, and 3 by all three. There is no shared canonical set of pages that answer engines agree on, at least not at this basket size.
One honest note about what this can and cannot tell you: our most-cited page is a roundup that names agencies, our own included. We have written before about the mechanism, which is that engines reading a category question tend to reach for lists rather than for the most on-topic single page. This study measures that concentration. It does not establish why, and we are not going to dress up a plausible story as a result.
Our own site against the same standard
The second half of this study is uncomfortable to publish, which is most of the reason to publish it.
We scored every URL in our own sitemap against a six-criterion answer-engine-readiness standard, at two commits six days apart. In between, we deliberately ran a formatting retrofit across the site. Both snapshots were built locally and scored by the same program, run unchanged, so any difference between the columns is a difference in the pages rather than in the ruler.
| Criterion | 2026-08-03 | 2026-08-09 | Change |
|---|---|---|---|
| C1 Answer-first lede | 97.4% | 97.4% | 0.0 |
| C2 Extractable facts with sources | 37.0% | 36.6% | -0.4 |
| C3 Question-shaped headings | 76.9% | 94.8% | +17.9 |
| C4 FAQ schema matching the page | 74.2% | 90.8% | +16.6 |
| C5 Visible freshness signal | 10.4% | 91.8% | +81.4 |
| C6 Claims integrity | 32.6% | 33.1% | +0.5 |
The sitewide average moved from 3.28 to 4.45 out of 6. Pages clearing 4 or more went from 556 to 1,080, which is 93.3% of the site. Pages clearing all six went from 1 to 2.
Every criterion that moved is a structural one, and every criterion that did not move is an evidential one. Question headings, FAQ parity and a rendered review date are template work. One engineer can change them across a thousand pages in an afternoon, and we did. Sourcing a figure is not template work. It is somebody reading a claim, finding out whether it is true, and either citing where it came from or deleting it. That is why 792 of our 1,158 pages still link to zero external sources, and why 66.9% of the site still fails claims integrity after a week of concentrated effort.
We are stating this about our own site and we think it generalises to the shape of the problem rather than to our specific numbers. Any site can be made to look answer-engine-ready in a week. Making it actually worth quoting is a different kind of work on a different timescale.
Does the score predict citation? On our data, no.
The obvious next question is whether the pages that score well are the pages that get cited. We joined the two datasets to find out, and the answer on this site is that the score does not separate them at all.
The 22 cited pages average 4.14 out of 6. The 1,136 uncited pages average 4.45. The cited pages score slightly worse, and none of them scores 6 of 6. Their individual scores run from 2 to 5.
There is a structural reason this comparison has little power left, and it is our own doing. After the retrofit, 93.3% of the site clears 4 of 6. A standard that almost every page passes cannot explain why 1.9% of pages get cited. We built a ruler that no longer discriminates, which is worth knowing before anyone treats a score like this as a citation forecast, ours included.
Two criterion-level differences point in interesting directions and neither one clears significance:
- Cited pages pass extractable facts with sources at 50.0% (11 of 22) against 36.4% sitewide. Fisher exact two-sided p = 0.19.
- Cited pages pass claims integrity at 18.2% (4 of 22) against 33.4% sitewide. Fisher exact two-sided p = 0.17.
This is where we have to correct ourselves in public. Our own internal audit on 2026-08-03 reported the second of those as a finding: that answer engines preferentially quote our least-sourced pages, because unsourced pages are where the dense numbers live. It is a good story and we believed it. On a fresh, independent probe with a significance test attached, it does not clear the bar. The direction repeats. The evidence does not support calling it a finding, so we are not calling it one, and we would rather say that here than let the earlier version keep circulating.
What this measurement cannot see
Two things we went looking for and could not get, both of which are routinely asserted by people selling in this category.
AI referrals carry no query. We wanted to connect citation to real visits, so we went through our own analytics for the 28 days from 2026-07-12 to 2026-08-08. There were 235 AI-referred sessions across six sources. The referrer on those sessions is a bare hostname with no path and no parameters, and the search-term dimensions are unset on effectively all of them. We do not know what a single one of those visitors asked, and neither does anybody else looking at their own site analytics. Every question class in this study is one we chose in advance, not one we observed a visitor use.
We are not publishing a share-of-traffic figure. The obvious next line would be "AI is X% of our traffic". We can count 235 AI sessions confidently because the source is explicit in the referrer. We cannot state the denominator with the same confidence, because direct traffic on this property is heavily contaminated by automated hits we have not finished excluding, and a percentage built on a denominator we distrust is a number that reads as precise and is not. So the count is published and the share is not.
We would rather leave two obvious holes in the write-up than fill them with figures we cannot defend.
If you want this run against your own site rather than ours, that is the work we do. Buzzbox Media is a Nashville agency working exclusively in medical device and healthcare marketing, and this measurement is the front half of how we approach answer-engine visibility for a client.
How this study was built
What we set out to do. Establish, for one medtech site, how often AI answer engines cite it, whether that result is stable enough to manage against, and whether a published readiness standard predicts it.
The citation probe
The questions. 13, fixed in advance. Five are agency-recommendation questions of the "best medical device marketing agency" shape. Eight are informational questions about medtech marketing practice. They are the same basket our standing internal tracker uses, so this run stays comparable to earlier baselines rather than being a new instrument measuring a new thing. Every question is printed verbatim in the dataset. The basket is a sample of question shapes we chose, not a census of demand, and a different basket would produce different rates.
The engines. Three, each through its vendor API with the web-search tool enabled: ChatGPT on gpt-4o with web_search_preview, Claude on claude-sonnet-4-5-20250929 with web_search_20250305, and Grok on grok-4-fast with web_search. The exact model string is recorded on every row, because "ChatGPT said" is not a reproducible claim and "gpt-4o with web_search_preview on 2026-08-09 said" is closer to one. Google Gemini is deliberately absent. Its grounding metadata returns redirect wrappers rather than resolvable publisher URLs, so we cannot count page-level citation for it on the same basis as the other three, and including it on a different basis would have made the totals incomparable.
The prompt. One wording, held identical across every engine and every run: the question, followed by "Cite the specific web sources you used." We did not tune per engine, and we did not ask the engine to consider us.
The runs. Five independent calls per question per engine, issued on 2026-08-09, giving 13 x 3 x 5 = 195 calls. No conversation state carries between calls. All 195 returned a full answer and none failed, so the dataset is a complete census of the run.
How citation was counted. Two outcomes, recorded separately on every call.
- Cited as a source is TRUE when a buzzboxmedia.com URL appears in the structured source list the API returns, meaning ChatGPT's message annotations, Claude's text-block citations and web-search tool results, or Grok's annotations and citations array. It is never regexed out of the answer prose.
- Named in the prose is TRUE when the token "buzzbox" appears in the answer text after removing markdown link targets, angle-bracketed URLs, bare URLs and bare host tokens, in that order. "Buzzbox Media is a Nashville agency" counts. "(source: buzzboxmedia.com)" does not.
A page counts once per call however many times an engine cites it within that call, so the 156 page citation events are call-level, not mention-level. The full raw response for every call was captured before any of these fields were computed, which is what let us find and repair the naming-measurement error described above without re-running the API.
The self-audit
The page universe is the sitemap itself, not a crawl and not a sample. 1,159 URLs, of which 1,158 resolve to a built page in both snapshots and are scored. The one exception has no built page and is excluded rather than counted as a failure.
The two snapshots are commits 088d174a and 6e7bcd8a, six days apart. Each was checked out into its own isolated worktree, built locally, and scored from the resulting static HTML with script, style, noscript, SVG, iframe and form content removed, so what is scored is what a reader sees.
The scorer is one program run unchanged against both builds. This matters more than it sounds: a before-and-after where the ruler also changed is not a before-and-after. The one date constant inside it, used for the 180-day freshness window, was set to the same value for both runs.
The criteria are the six defined in the dataset dictionary below. Two properties of the standard are worth stating because they change what passes. FAQ schema only passes when its questions and answers actually appear in the visible page, so schema asserting content the page does not carry fails. Freshness only passes when the date is both real and rendered, so a correct dateModified that no reader can see does not count.
Statistical tests. Proportions carry Wilson score intervals. The two cited-versus-uncited comparisons carry Fisher exact two-sided tests, computed on the raw counts, and are reported with their p values whether or not they clear a threshold.
When this was last reviewed, and why that matters here more than usual. Reviewed 2026-08-09, the same day the probe ran. Engine behaviour is versioned and undocumented, so a citation rate ages in a way a conference price does not: a number from this page read a year from now describes three model versions that may no longer exist. We re-run the probe rather than leave a stale measurement reading as current, and the review date above moves only when the underlying run does.
Limitations, stated plainly
- This is one site. It is the largest limitation by a wide margin. Every rate here describes buzzboxmedia.com and nothing else. We have no basis to say what an AI citation rate looks like for medtech sites generally, and a single-site study cannot produce one. Treat the method as the transferable part and the numbers as ours.
- We are the publisher and the subject. We chose the questions, we built the instrument, we scored our own pages, and a favourable result would be commercially useful to us. That is a real conflict and it is why both datasets ship in full, why the naive measure ships alongside the corrected one, and why the findings that make us look worse are in the key findings rather than in a footnote.
- Five runs is a coarse instrument for stability. With 5 runs, a pair whose true per-run probability is 0.5 appears perfectly stable 6.25% of the time, and one at 0.8 appears to be a clean sweep 32.8% of the time. Our stability counts therefore understate instability. They do not overstate it, which is the direction that matters for the claim we are making.
- The 13-question basket is a sample of question shapes, not of demand. We picked questions our pages are built to answer. A basket chosen adversarially would show lower rates and a basket chosen generously would show higher ones. The questions are published so this can be judged rather than assumed.
- Engine behaviour is versioned and undocumented. These are three specific models on one day, with web-search tools whose retrieval behaviour the vendors do not publish and can change without notice. Nothing here should be read as a durable property of "ChatGPT" or of AI search in general.
- Gemini is excluded for a technical reason, not a judgment about it. Its grounding URLs are redirect wrappers we cannot resolve to publisher paths, so page-level citation is not measurable for it on the same basis. Its absence means "engines" in this study means three engines.
- Citation is measured through APIs, not through the consumer products. The web-search tool an API exposes and the retrieval stack behind a consumer chat interface are not guaranteed to be the same system. We measured what we can measure reproducibly and cannot verify that a person typing the same question into the app would see the same sources.
- The cited-versus-uncited comparison rests on 22 pages. Neither criterion-level difference reaches significance and we report both p values rather than the one that suits us. Nobody should act on either as if it were established, and that includes us.
- The self-audit standard is our own. The six criteria are a house standard, not an external or agreed one. Another analyst would define them differently and get different pass rates. The per-page results ship so the standard can be argued with, and after the retrofit it no longer discriminates well enough to be used as a predictor.
- We cannot connect any of this to revenue, and we did not try. The citation data and the analytics data cannot be joined, because AI referrals carry no query. Any claim linking a citation to a visit, a lead or a deal would be an inference dressed as a measurement.
Conflict of interest. Buzzbox Media is a marketing agency that sells answer-engine visibility work to medtech companies. The site measured is our own. No client data of any kind was used, no client site was probed, and no third party reviewed this before publication.
Reuse. Both datasets are free to download and republish with attribution to Buzzbox Media and a link back to this page, under a CC BY 4.0 license. Suggested citation: Buzzbox Media, AI Answer Engine Citation Study, 2026. Two asks. Please do not re-sort the third-party host column into a ranking of which companies AI engines prefer: it records what happened to be in a source list on one day and is not a standing measure of anyone. And please carry the run count alongside any citation rate you quote from this, because a rate with no runs behind it is the thing this study exists to argue against.
Download the datasets
Both files are published in full, with no gate and no email required. Every figure on this page recomputes from them.
Every engine call made for this study, one row each. 15 columns, 195 rows.
Our own site scored against the six criteria at both commits. 24 columns, 1,158 rows.
The probe file carries one row per engine call, including the calls where no page of ours was cited, because a citation rate whose denominator is hidden cannot be checked. It also ships a column called brand_string_anywhere holding the naive brand test we started with and rejected, so the methodology correction described above can be verified rather than taken on trust.
The page-score file carries every URL in the sitemap, including the 792 pages that link to zero external sources and the 775 that fail claims integrity. Publishing the pages we fail on is the point of a self-audit. Rows are sorted by URL, deliberately not by score.
What we plan to do next, on the record
Three things, so this can be held to.
- More runs, fewer conclusions. Five runs bounds instability loosely. The next run raises the run count before it raises the question count, because the stability estimate is the weakest number here and it is also the most important one.
- A second site. Everything above is n = 1. The finding worth having is whether run-to-run instability looks like this on sites that are not ours.
- The evidence half. We have 792 pages linking to zero external sources. We are going to work on that number and report it again against this same baseline, whichever direction it goes.
If you want the execution side of this rather than the measurement side, our medical device SEO checklist covers the on-page work, and our other published dataset applies the same standard to medical conference booth pricing.