Semrush spent four months running 126 million real AI search prompts through ChatGPT, Gemini, Google AI Mode and Google AI Overviews, then published the results as its 2026 AI Visibility Index. The headline number wasn't a brand ranking. It was this: only 45% of marketing leaders can accurately measure how their brand shows up in AI-generated answers, and just 9% have a tool that tracks it across more than one platform. That gap has a name now. Prompt volume.
Prompt volume is the count and pattern of real user queries that pull a brand into an AI-generated answer, tracked at scale instead of guessed at from a handful of test questions typed into ChatGPT the night before a board meeting. It works the way search volume worked for Google, except the query never touches your analytics, and the "result" is a paragraph the model wrote instead of a link someone clicked. Teams that used to eyeball ten sample prompts are now expected to show a documented set running into the thousands, refreshed on a schedule, broken out by engine.
That expectation didn't come from nowhere. Semrush's 126-million-prompt sample broke visibility down by sector, and the spread was wide. In News and Media, the top three brands captured 82.9% of all visibility inside AI answers. In Finance, the top three held just 41.4%. A finance brand chasing "top three" the way a media brand does is solving the wrong problem, because the opportunity in a fragmented sector looks nothing like the opportunity in a concentrated one. Buyers now ask vendors to show sector-adjusted numbers before they sign, not a flat visibility score pulled from someone else's industry.
What Prompt Volume Actually Measures, and What It Doesn't
Prompt volume by itself only counts how often a query surfaces your brand somewhere in the answer. It says nothing about whether you were named favorably, cited with a link, or buried under four competitors in the same paragraph. Semrush's data makes that distinction hard to ignore: ChatGPT cites an average of 15 sources per response, Gemini cites roughly 3, and on Gemini the brands mentioned in the answer text and the domains actually cited overlap only 30% of the time. A brand can be named constantly and still be invisible in the citation layer. The reverse happens too.
Treat the raw number as a filter, not a scoreboard. It tells you which prompts are worth digging into further. The digging is where mention rate, citation rate, and real LLM brand visibility benchmarking against named competitors actually live.
Building a Prompt Set That Reflects Real Buyers
Start from language, not keywords
Most teams build their first prompt set the way they built keyword lists a decade ago. Guess ten phrases, run them once, screenshot the results for a slide. That produces a vanity number and nothing repeatable. A usable prompt set starts from the language buyers actually use when they ask an AI model for a recommendation, not the head terms a brand wishes it ranked for. Pull real questions from sales call transcripts, support tickets, and the forums where your category gets discussed. Then sort by buying stage. Awareness prompts, comparison prompts, decision prompts. Each behaves differently across engines and needs its own tracking line.
Track the number by engine, not as one blended score
A blended visibility score hides more than it shows, because ChatGPT, Gemini, Google AI Mode, AI Overviews, Perplexity and Claude pull from different indexes and cite in different ways. Semrush found only 36 global brands held top-100 visibility across all four platforms it studied, every month, for the full period. Most brands that look strong on a single combined number are actually strong on one engine and close to absent on the rest. Report each engine on its own first. Blend only once you know which one is carrying the average.
Re-run it on a schedule, because the answer moves
A single snapshot of prompt volume is out of date by the time it reaches a slide deck. Model updates, retrieval changes and competitor content all shift which prompts surface your brand, and Semrush's own January-to-April comparison showed real movement inside a single quarter. Monthly re-runs, at minimum, are what separate an actual LLM visibility program from a one-time audit somebody ran to answer a question from the board.
Where Teams Get This Wrong
The same mistakes show up in almost every prompt volume rollout. Treating one run as truth. Comparing raw counts across sectors with wildly different concentration instead of a sector-adjusted benchmark. Reporting mentions while ignoring the citation layer, which is where a recommendation is actually earned. Fix the reporting cadence and the sector context first. The tooling matters less than the discipline behind it.
What To Do With This Number Monday Morning
Pull ten real buyer questions from your last twenty sales calls. Run them through ChatGPT, Gemini and Google AI Mode this week, by hand if that's what it takes. Note who gets named, who gets cited, and where your brand is missing entirely. That is your baseline. Everything Semrush's 126-million-prompt study found applies just as much to ten prompts as it does to ten thousand.
Prompt volume isn't a number you buy a tool for and forget about. It moves every time a model updates, and the brands checking it monthly are the ones showing up when their buyers ask an AI which vendor to pick.
Want the sector benchmarks, the engine-by-engine breakdowns, and what's moving in ChatGPT, Gemini and Perplexity before it shows up anywhere else? Subscribe to the newsletter below.


