In September and October, HubSpot's numbers went the wrong way. Across the questions they track, ChatGPT was recommending their brand less often than it had the month before. If you were watching a single visibility score, that is the moment you call an emergency meeting and start unpicking whatever you shipped in August.
Nothing they shipped in August was the cause. Aja Frost, senior director of global growth at HubSpot, described the episode at GROW Europe 2025: ChatGPT had turned the dial down on brand visibility generally, and brands across the board were being recommended less (HubSpot Live). The only reason they knew that, rather than blaming themselves, is that they were watching a second metric alongside the first.
Why can't you judge AI visibility from one number?
Because a single visibility score cannot distinguish between your performance changing and the platform changing underneath everyone. Frost's September and October drop is the cleanest illustration available: absolute visibility fell, share of voice against competitors held, and only the second measurement revealed that the cause was industry wide rather than self inflicted. Any measurement setup with one number in it will produce confident wrong conclusions roughly as often as right ones.
Key takeaways
- Frost's scorecard has four metrics: AI visibility, AI share of voice, AI citations, and AI referral demand. Each answers a different question and none of them is sufficient alone.
- Visibility tells you whether you are recommended. Share of voice tells you whether a change was yours or the platform's.
- Citations matter beyond the count, because HubSpot reports being described more favourably when its own content is the cited source.
- Referral demand is the hardest to attribute, and Frost's recommendation is to ask buyers directly rather than trust analytics.
- The same prompt returns different sources on different runs, so any score derived from a single query is overstating its precision, including one from our tool.
What are the four things actually worth measuring?
Frost's scorecard covers visibility, share of voice, citations and referral demand, in roughly that order of directness. Visibility asks whether you are recommended for the questions you care about, on each engine, and she calls it the north star. The other three exist because a north star on its own does not tell you where you are standing.
Working through them in order is a reasonable way to build a measurement routine, since each one costs more effort than the last and answers a question the previous one raised.
How do you measure AI visibility?
Pick the questions your buyers actually ask, run them on each engine that matters to you, and record whether your business is named in the response. That is it, and its simplicity is the reason Frost treats it as the primary metric: it corresponds directly to the outcome you want, which is being one of the handful of businesses an assistant mentions.
The discipline is in the setup rather than the running. The question set has to be the questions buyers type rather than your brand name, because searching your own name tells the system the answer. The session has to be neutral, since Frost notes that responses are shaped by prior questions, previous clicks and connected accounts such as email and CRM, which means checking your own visibility while logged into your own account produces a flattering and useless result. And the question set has to stay fixed over time, because a metric that changes its own definition month to month measures nothing.
What is share of voice, and why does it matter more than it sounds?
Share of voice asks a comparative question: of the responses that recommend any solution at all, how often is the one recommended you rather than a competitor. It converts an absolute number into a relative one, and the September and October episode is the argument for it.
Frost's account is worth restating carefully because it generalises. Visibility fell. Read alone, that says the recent work failed and something needs undoing. Share of voice held steady, which said the pool of recommendations had shrunk for everyone and the relative position was intact. The correct response was to do nothing, which is exactly the response a single number would never have produced.
The same logic runs in the other direction and is easier to miss. Visibility can hold flat while share of voice falls, because a platform got more generous with recommendations generally and competitors captured the extra room. That looks like stability and is quietly a loss.
For a small business this is the metric that most changes what you do with a report. The useful output of a visibility check is rarely your own score. It is the list of who does get named instead of you, because that list is the actual competitive set as an AI understands it, which is often not the one you had in mind.
How do you measure citations, and why track them separately?
Citations count how often an engine uses your own content as the source for what it says, which is different from being mentioned in the answer. You can be recommended without being cited, when the engine draws its facts from a review site or a directory and names you as a conclusion, and you can be cited without being recommended, when your explainer supplies the background for an answer that then recommends somebody else.
Frost's reason for tracking them separately is the more interesting part. HubSpot reports that when their own content is the cited source, their brand tends to be described more positively and placed higher in the response. That is HubSpot's own observation about its own results rather than published research, so treat it as a hypothesis worth testing rather than an established mechanism. It is a plausible one: the description of you in an answer comes from whatever source the engine trusted, and your own page is the source most likely to describe you the way you would.
This is also where mentions and citations get conflated in vendor reporting, frequently in whichever direction makes a chart look better. If you are buying a tool, ask which one it counts.
How do you measure demand you cannot attribute?
You ask. Frost is direct that AI referral demand is hard to attribute and recommends a post purchase survey asking how the buyer heard about you, rather than trusting analytics to catch it.
That is a low technology answer to a high technology problem and it is the right one. A buyer who read three AI answers, formed a shortlist, then typed your company name into a browser arrives in your analytics as direct or branded search. Every step that created the demand is invisible, and the step that merely captured it takes the credit. No amount of tagging fixes this, because the events happened somewhere you have no instrumentation. The only reliable sensor is the buyer, and the only reliable time to ask is after they have bought, when the answer is a memory rather than a guess.
The trend data supports bothering. Frost told GROW Europe that more than half of buyers now say they use AI to make purchasing decisions, and that AI Overviews appear in nearly six in ten searches and halve click through rate when they do.
Why does the same prompt give different answers?
Because these systems are probabilistic, and this is the limit that constrains every measurement above. Visibility varies across paraphrasing, interface, model version, location and whether you are logged in (Search Engine Land), and an academic study of the problem found measurement error only falls to an acceptable level after roughly seven repeated runs of the same prompt (arXiv:2604.07585).
Put that next to the September and October episode and the picture is uncomfortable in a useful way. Your reading moves for at least three reasons that have nothing to do with you: run to run randomness, platform level policy changes, and personalisation from whatever session you happened to use. Only the fourth reason is your own performance, and it is the smallest of the four in any short window.
There is a further problem with no fix currently available. There is no industry standard for what an AI visibility score means. Different tools use different prompt sets, weight engines differently and define a mention differently, so two vendors scoring the same business will disagree and neither is checkable against the other.
What does an honest measurement routine look like?
Fewer numbers, held to a stricter method, read over a longer window. The practical version for a business that is not going to hire an analyst:
- Fix a question set of five to ten questions your buyers genuinely ask, phrased naturally and without your brand name in them. Never change it, or you lose the comparison.
- Run each question several times, from a neutral session with no memory or location bias, on at least two engines. One run is an anecdote.
- Record who else appears, every time. The competitor list is the share of voice signal and it is usually the most actionable line in the whole exercise.
- Re measure monthly, not weekly. More frequent checking mostly measures the noise described above.
- Ask every new customer how they first heard about you, and keep the answers where you can count them.
- Read direction, not level. Whether you moved and in which company, over three months, is a real signal. A precise score in a single month is not.
SignalCheck covers the first two of those steps, running a live query from a neutral session with no saved memory or location and showing the raw answer text rather than only a score. That last part is deliberate, and it is where our own honesty note belongs: a score produced from a small number of queries on one day is directional at best, and if we reported it to two decimal places we would be claiming a precision the underlying systems do not support. Read the answer text, note who got named, and repeat it next month. On how long changes take to show up, see how long AEO takes to work.
What about the outcome numbers vendors quote?
Read the source attached to each one, because the gap between measured research and self reported marketing is wide in this category. Frost reported HubSpot's own results since launching their AEO strategy as the highest share of voice in their category, citations improved by 433 percent, and demand up by nearly 2,000 percent. Those are HubSpot's own reported figures about HubSpot, with no control group and no published methodology, from a company that also sells tooling in this space.
That is not a reason to dismiss them. It is a reason to file them as directional claims from an interested party rather than as evidence of what your business should expect. The same test applies to every percentage in a GEO vendor's case study, and to ours. The measured, independent data in this field, from Pew, Ahrefs, Cloudflare and Adobe, is a much smaller set of numbers and a much duller one, which is generally what independent data looks like.
Frequently asked questions
How many prompts should I track? Five to ten real buyer questions is enough for most small businesses, and it beats a hundred that nobody actually asks. What matters far more than the count is that the set stays identical between measurements.
Should I track every AI engine? Track the ones your buyers use, which for most businesses means starting with ChatGPT and Google AI Overviews and adding others only if you have reason to. Coverage across many engines is a vendor selling point more often than it is a business need.
My visibility dropped this month. What should I check first? Whether your competitors dropped too. If the whole set moved, it is likely a platform level change of the sort Frost described in September and October, and undoing your own recent work would be the wrong response. If only you moved, then look at your own changes.
Is a visibility score from a paid tool more reliable than checking manually? It is more consistent, which is not the same thing. A tool runs the same prompts the same way every time, which makes the comparison over months meaningful. It does not have access to any ground truth about the engines that you lack, and no vendor has solved the run to run variance.
Can I just look at my analytics for AI referral traffic? Only partially, and it will understate the effect substantially. Buyers routinely research in an AI tool and then arrive by typing your name, which records as direct or branded traffic. This is why Frost recommends asking buyers directly instead.