Second order thinking
Stop calling it "AI visibility"
It promises a measurement nobody can make. Here is what you can actually see
"AI visibility" is fast becoming a phrase that means everything and nothing.
Ask three marketers what it means, and you will get three different answers. They might include showing up in Google's AI Overviews; "mentions" in ChatGPT; an AI visibility score from an SEO tool; or a share-of-brand-voice percentage on a dashboard. All of them are real numbers; each measures something different, but they are all filed under two words that let everyone nod along but help nobody take action.
That confusion makes it hard to understand the impact of AI on your marketing activities.
And there's a second, more important issue hiding inside it: "visibility" promises something that can't actually be delivered. No tool you can sign up for can accurately count how often an organisation appears inside the answers buyers see. That would mean interrogating the models at scale and trusting whatever they happen to say back.
Even the AI operators, who see the answers directly, publish almost nothing: Google alone reports how often your pages appeared, but not what the answers said about you, and the others report nothing at all. So the promise at the centre of "AI visibility" is one that nobody can keep.
Here is the term we use instead: citability. It is deliberately one step back from visibility. This is not about whether you are "in the answer", which you can't measure; instead, it asks, "Are you the kind of source the machine can and does reach for?"
Are you readable, structured, authoritative and trusted enough to be cited, and are the large language models actually reading your content and pulling it in to answer real questions? That is activity you can see. Citability is what everyone is implying when they say "AI visibility", plus the proof from real data.
"AI visibility" hides three problems that need three different responses:
AI Overviews · AI Mode · Gemini
Same crawl, same index as search
Observe via GoogleChatGPT · Claude · Perplexity
Their own bots read your content live
Observe directlyWhat assistants say unprompted
No fetch: answers drawn from training
Interrogate onlyThe first is Google's AI answers, and they are the most misunderstood:
They are not a separate AI system bolted onto Google. All of them run on the same Googlebot crawl, the same index and the same ranking that power ordinary search, with the Gemini models writing the answer on top. Even the Gemini assistant has no crawler of its own; it reads your pages through the Search index. There is no separate "AI Googlebot" to win over.
Your visibility in Google's AI answers and your SEO are the same job, because they run on the same index.
The second is the genuinely new thing: the answer engines. ChatGPT, Claude and Perplexity crawl your site, read your content, and cite it in their answers independently of Google, on their own infrastructure, with their own bots. This is the new audience on your site, and it behaves nothing like a search engine. It is an intermediary: it sits between your expertise and your buyer's question, and it decides what gets said about you. This is where citability lives.
The third is the least obvious: what the models say about you from memory. When an assistant answers without fetching anything, it describes you based on whatever its training runs ingested months ago. That reading happened in traffic only you could see, in your server logs; to every outside tool it is invisible. And the answers built on it can't be observed by anyone: you can only interrogate them by asking the assistants what they say about you and treating every answer as a snapshot rather than a score.
Three different problems: one is the same job as your SEO, one you can observe directly, and one you can only interrogate. Calling all three "AI visibility" guarantees you treat all three like SEO.
A whole industry has emerged, and it has named the new work: GEO (generative engine optimisation) and AEO (answer engine optimisation) are the modern-day counterparts to the old world of SEO. If you read the playbooks, you'll find advice similar to that of the last fifteen years, with "for AI" added at the end. Add FAQ schema. Write clearer headings. Earn more mentions. Almost all of the technical advice is sensible. Almost none of it is measured.
That is the real issue, and it is a mindset problem rather than one created by the vendors. It is SEO thinking transplanted onto a new technology that nobody has actually observed. Do these ten things and you might appear in ChatGPT. The big assumption there is "might". There is no reliable data on the other side to tell you whether any of it worked, because the vendors selling the tactics cannot see what the LLM did with your content.
Stop treating AI like a channel
Here is a test that cuts through the whole debate: ask who can actually count your appearances. Only the operator (Google, OpenAI, Anthropic, etc.) can, because only the operator sees the answers it gives. And as we said at the top, only one of them does.
Google's Search Console generative AI report shows how often your pages appeared in AI Overviews and AI Mode (Google counts these as "impressions"). It is first-party, it is real, and we track it across every site we monitor. Notice its limits, though: even Google, which owns both ends of the exchange, tells you how often you appeared, but not why.
So any tool that claims to score your visibility in ChatGPT is measuring a system whose owner provides no way to measure it. It cannot count. It can only ask sample questions (many times) and extrapolate, which is a guess with a dashboard in front of it.
And even a perfect count would be a crude one. A mention is not a recommendation: the same score covers the answer that names you as the obvious choice and the answer that names you as the one to avoid. Counting appearances, even with numbers attached, tells you that you came up. It cannot tell you how you came across.
This stopped being hypothetical for us this summer. A client was handed a vendor's AI-visibility report built the standard way: run a sample of prompts against the models, note who gets named, score the site. The verdict was a 3 per cent citation rate across four key pages, dressed up as evidence that the site's foundations were poor.
Then we looked at the server logs for those same four pages. In five and a half weeks, the AI engines read them nearly 300 times. Nearly a quarter of those were ChatGPT live-reads. That's when a human asks a question, and the machine pulls that exact page in real time to answer it. On one of the four pages, the machines actually out-read the human visitors, 1.6:1.
The sampled guess in the report said these pages were invisible. The server logs said the engines reach for those exact pages dozens of times a month, all verified against the operators' published IP ranges to ensure accuracy.
So the problem is not GEO or AEO, and it is not any one vendor. The problem is working blind: optimising without evidence, scoring without measurement, guessing when the machines leave a record of every visit in your server logs.
You don't have to guess. You can observe.
Since June, we have been reading the server logs of more than a dozen client websites. We have analysed more than three million hits, identified the AI crawlers and what they read, and checked each one against the LLM operators' published server IP lists.
So when we say OpenAI's assistant fetched pages on the busiest site roughly 24,000 times in July to answer real questions, it is not a guess from a user-agent string. We checked the IPs because user-agents can be faked, and we would rather be right than overstate the numbers. July was our first clean, full calendar month across all sites, and it is the basis for all data below.
One caveat before the findings: the July figures below come from eleven sites and two clean months, a relatively small base. It is enough to observe the patterns below, which repeat across every site in the set. It is too early to establish if there are any trends. Nothing here is a benchmark yet, but we hope to build this over time and report back when we have more data.
Three things show up the moment you stop guessing and start measuring:
Across the sites in July, three in every five bot requests came from AI systems: training crawlers, AI search indexers and live assistants, around the clock rather than in daytime peaks. The mental model of "Google plus a few SEO tools" is no longer a valid view of your bot traffic.
Across the portfolio in July: more than 70,000 genuine AI reads led to fewer than 500 human visits arriving from an AI answer. About 140 reads for every visit. Judge that by normal channel measurements, and a CTR of less than 1% looks insignificant. But it's not like a normal channel. It is the machine reading your content at the moment of the question, and 140 to 1 is the reach, not a disappointment.
July, across the client portfolio: more than 70,000 genuine AI reads, fewer than 500 human visits arriving from an AI answer
On one client site, nearly four in five live reads were insight content. On another, the machines read the news and knowledge section eleven times for every human page view, while the product pages drew fewer machine reads than human views. The pages built for human conversion are not the pages the machines reach for. Your content strategy and your citability strategy are the same thing.
The strongest signal in there is the live read: ChatGPT-User pulling a specific page because a real person just asked about your organisation or what you do in the chat. You can't see the answer it gave, but you can see it read the site. That is observed citability, not a guessed visibility score.
And there is now a second signal, on the other side of the question: the first humans arriving from an AI answer. When someone clicks through from chatgpt.com or perplexity.ai, the machine names you in an answer, and they trust you enough to click. Every one of those is proof a citation happened. It is a baseline, not a count, because most citations never produce a click.
GA4 now shows these in the channel reports under "AI Assistant", and it is worth checking that out. But GA4 is consent-gated, and we can now put a figure on the gap: when we switched on consent-independent referral counting from one client's server logs, GA4 turned out to have been undercounting referrals by 50%. The logs capture every arrival, consent or not, which is why we track the trend there.
So we measure the real data:
In between those sits the one thing no one outside Google can see: what the person asked, and what the answer actually said about you. That gap is why Google's count and your server logs belong together. Google tells you how often you appeared, but nothing about the answer itself. The logs tell you what the machines are reading, and what they read is the raw material of every answer they give about you.
Appearances tell you that it happened. The logs tell you why.
Every figure in the last section is a snapshot, and the data is constantly changing. The value lies in the trend, and a trend becomes visible only if you measure consistently over time.
The Search Console AI data and the server logs confirm the volatility. A single article page on one client site appeared in Google's AI results roughly 4,600 times in June and roughly 32,700 times in July. That is a sevenfold move in a month, on one page, with no change to the content. Anyone managing that as a static score is working with out-of-date data.
July also showed us the other direction. For one client, Perplexity was the second-most active bot in June, at 23 per cent of their citability. In July it fell to 8 per cent, with daily reads down about 80 per cent. Perplexity publishes its server list, and 98 per cent of the traffic over both months was validated as real. One of the engines changed its behaviour. We don't know why, or whether it will return, but we are building a picture of each AI operator's habits over time.
A prompt-tracking dashboard would have kept showing the same brands being mentioned while the site activity collapsed by four-fifths. A score taken in June was wrong by July. Observing caught it inside a month, put a size on it, and turned a vague worry about AI concentration into a specific, answerable question: what do we do about the fact that one engine now carries nearly nine-tenths of our machine readership?
That is the argument for continuous observation rather than a single score. A score you check once a quarter is last quarter's number; the movement in between is where the insight is.
A fair question: how does our own site do? We are in the foothills, and we would rather say so than pretend otherwise. But in July, the machines read our site nearly six times for every human view. Before we get carried away, ours is a small site and a small month, so we are treating that ratio as directional rather than definitive.
And the detail is humbling: the single biggest visitor to our site is still a training crawler, Anthropic's ClaudeBot, which read us roughly twice as often as every live assistant and AI search bot combined. Claude is learning about us, but not yet citing us. Meanwhile, GA4 recorded no humans arriving from an AI answer all month. The irony is that the pages the LLMs read most on our site are the ones we wrote about citability. Our argument is being ingested. We now need it cited.
Nearly six pages read by machines for every one a human opened. Small site, small month: read as direction, not a definitive ratio.
We said in June we would publish our own performance as it moves, up or down. It moved up: our AI reads rose roughly 38 per cent month on month, nearer 12 per cent once you correct for June being a short measurement window. We give you both because one flatters us and the other is true to the day count, and you should be suspicious of anyone who shows you only the first kind.
Our profile follows the same pattern we see across the knowledge-based organisations in our client base, and is heavily skewed towards our "Second order thinking" article content. So you will see more of that over the next few weeks as we expand on our findings.
This is not pedantry about terminology. The words you use determine what you can act on, and what you can honestly claim.
"AI visibility" keeps the problem comfortably vague and promises a measurement nobody can make. "Citability" names something specific and one step earlier, the part you can actually see. It separates the new work from the SEO you already do and from the Google work that runs on the same index, and it points at evidence you can put in front of a board this quarter rather than a debate you can have forever.
And beneath the terminology lies the real reframing. The reason "AI visibility" leads everyone into the optimisation trap is that it treats machines as a channel: a funnel you can optimise and judge by what comes out at the end.
The numbers above describe something else. An audience. It arrives daily, it chooses your thinking over your service pages, it changes its behaviour from one month to the next, and it reads you 140 times for every visitor it sends.
That is not a channel to optimise. It is an audience to understand, and citability is simply what you measure when you start treating it like one. Mapping that audience, who is actually in it, what each of them reads, and how far each has travelled towards citing you, is the subject of our next article.
The question is not "is AI relevant to us". It is "when a buyer asks the machine, are we relevant to AI?
Are we the source it reaches for, and can we prove it?"
If you want a quick view of where you stand today, the Perception Gap Scorecard takes a few minutes and needs no conversation with us.
If you want to know what one audience actually sees when it meets your organisation, our free Second Opinion assesses your organisation from the outside, through two lenses: how the human audience finds and reads you, and how the machines find, read and represent you.
One audience of your choice. To ensure high quality, we only work with two organisations each month.
And if this article has raised a question your own server logs should answer, decoding them is where we would start.
Citability is whether you are the kind of source an AI engine can and does reach for: readable, structured, authoritative and trusted enough to be cited, with evidence that the large language models are actually reading your content and pulling it in to answer real questions. It is deliberately one step back from "AI visibility". You cannot count how often you appear inside AI answers, but you can observe, in your server logs, the machines reading your content at the moment of a question.
Mostly, no. Appearances inside AI answers can only be counted by the operator, and Google alone publishes anything: the Search Console generative AI report shows how often your pages appeared in AI Overviews and AI Mode (Google counts these as "impressions"), but not what the answers said about you. The other operators publish nothing. Tools that claim to score your visibility in ChatGPT sample prompts and extrapolate, which is a guess with a dashboard in front of it.
Observed behaviour rather than sampled guesses. Your server logs record every AI system that reads your content: training crawlers, AI search indexers, and live reads, where an assistant fetches a specific page because a real person just asked a question. Each one can be verified against the operators' published IP ranges. That is observed citability, not a guessed visibility score.
Yes, but far less than they read. Across the client websites we monitor, July saw more than 70,000 AI reads produce fewer than 500 human visits from AI answers: about 140 reads for every visit. GA4's "AI Assistant" channel shows these referrals, but it is consent-gated; on one client site it undercounted them by about half compared with the server logs. Judged as a traffic channel, AI looks insignificant. It is better understood as an audience reading your content at the moment of the question.
For Google's AI answers, SEO is the whole job. AI Overviews, AI Mode and the Gemini assistant all run on the same Googlebot crawl, index and ranking as ordinary search; there is no separate "AI Googlebot" to win over. The genuinely new work is with the answer engines, ChatGPT, Claude and Perplexity, which read your content with their own bots on their own infrastructure.
Second order thinking