A client asked me a question recently that I've been asked a lot this year: "How do you actually know which AI systems are reading our site?" Not why it matters, which Mike and I have covered elsewhere, but how. What are you looking at, and why should we trust it?
It's a fair question, because the answer most people get is a score. A tool puts a sample of questions to ChatGPT and the others, notes which organisations get named, and extrapolates. We've made the case against that elsewhere, so this article is about what we do instead: the method behind the monthly reports we produce for our clients and for our own site.
What is a server log and why does it matter?
Every website runs on a web server, and every time anything asks that server for a page, the server writes one line to a log file. That line records the time, which page was requested, whether the server was able to return it, the internet address (the IP address) the request came from, and a short piece of text called the user-agent, in which whoever made the request says what they are.
A person's browser sends a user-agent that says, "Chrome on a Mac". Google's crawler sends one that says "Googlebot". When ChatGPT fetches a page from your site to help answer someone's question, it sends one that says "ChatGPT-User".
That last part is the important bit; ChatGPT, Perplexity and Claude all fetch pages live when a person asks them something that your content might answer, and they identify themselves when they do. So your log file contains a record of every time an AI platform has reached for the content on your site.
It's important to note that Google Analytics doesn't see any of this bot activity. GA4 only counts a visit when a browser runs its tracking script and, in the UK, when the visitor has agreed to be tracked. A bot doesn't run scripts and doesn't consent, so it's invisible to "traditional" analytics.
What we're left with is two separate data sources that record different things:
- GA4 records the humans who landed on your website and agreed to be counted.
- The log records everything that touched the server: humans and bots, whether they consented or not.
When we say an AI assistant read a page 2,000 times and sent 15 people to the site, those are two separate counts from two separate sources, and it's the gap between the two numbers that's the finding.
Across the websites we monitor, the logs now record more than three million requests a month from automated visitors of one kind or another, and three in every five of those come from AI systems. On several of those sites, AI reading has overtaken human page views.
HOW WE DO IT
1 Get the logs somewhere you can work with them
The first thing to know about log files is that they're enormous! A busy professional-services site generates tens of thousands of log records a day. When we start looking at multiple sites, across the course of a full month, the count quickly ramps up to the millions.
Clearly, nobody's reading that by hand and it's well past what a spreadsheet will open. So before we can start counting anything, the logs have to be collected and put somewhere built for handling data sets of that scale.
For us that's BigQuery, Google's data warehouse:
- Our client sites send their logs to a log-management service in real time (for error monitoring).
- We built an automation that pulls the previous day's traffic from there into BigQuery.
- Once the lines are in, a piece of code reads each record and picks out the data we need.
Getting that working reliably took considerably longer than we expected.
The first hurdle is that log-management tools are built for looking up one request when something has gone wrong, not for handing over a whole day's traffic, and they cap what they'll return in one go.
The second hurdle is spotting errors when working with data on this scale. The first iteration of our data pipeline looked like it was working as intended; however, it was quietly returning only about a third of the records and reporting it as the whole. In the end we built dynamic rules to validate the volume of logs we were interested in analysing against the raw counts, to help flag any suspicious variance.
Once we have the data where we need it, the next step is working out what it means.
HOW WE DO IT
2 Identify what's actually reading your site
A month's worth of server log data contains hundreds of different user-agents, and the raw list isn't any use on its own.
Every one has to be matched both to the company operating it and to what it's doing on your site. That's less obvious than it sounds, because each operator runs several bots doing different jobs. We currently catalogue around sixty, and new ones appear most months.
What matters for reading the results from an AI citability perspective is the job each bot is doing, so we group them:
| Group | What it's doing on your site | Example |
|---|---|---|
| AI assistant, live | Fetching a page right now because a person has just asked a question | ChatGPT-User |
| AI search | Building the index an AI assistant looks things up in | OAI-SearchBot |
| AI training | Copying pages in bulk to train a model on | GPTBot |
| General bulk crawlers | Large-scale crawling that feeds AI products among other things | Amazonbot |
| Search engines, SEO tools, link previews | Everything else automated | Googlebot, AhrefsBot |
The key is the distinction between the first three:
- A live fetch means the assistant has reached for one of your pages at this moment, because a real person has just asked it something.
- An indexing crawl means it's cataloguing your pages so its assistant can search them later.
- A training crawl means an operator is copying your content to teach its model. That's about what the model might know in future.
The live group is the one to watch; ChatGPT-User fetching a page is the closest thing we have to being able to see the moment a question is asked.
There are two hurdles here too. The first is that the operators' own documentation is the only source for what each bot does, and it's uneven: some are clear, some are vague, and each operator's team of bots is different. Perplexity, for instance, runs no training crawler at all, so its profile on your site looks nothing like OpenAI's, and a report that compares them without knowing that will mislead you.
The second is the general bulk crawlers (Amazon's and Huawei's are the big ones), which are a large slice of the raw AI volume on most sites. Whether you count them changes the totals quite a lot, and a report should say what it's done with them.
With every line labelled, the structure of the data begins to look sensible, and the next step is to validate it.
HOW WE DO IT
3 Check the bots are who they say they are
This is the part we didn't fully appreciate when we started, and it's the part most likely to catch out anyone doing this for the first time.
A user-agent is just text, and anyone building a scraper bot can make it say whatever they like. 'ChatGPT-User' is a particular favourite, because a lot of sites have been told to let it through. So a raw count of AI user-agents is a strong signal, not a verified one, and a report built on the raw count alone is going to include imposters.
The way to check is the IP address, which is much harder to fake because it's where the request actually came from. In principle that's simple. In practice:
- Some operators publish the addresses their crawlers use, and every request can be checked against them.
- Some don't publish anything, which means their traffic can never be fully verified, however genuine it is.
- The published lists change, so a check that was right last month can be wrong this month.
Doing this properly means keeping those lists current and treating each request as verified, plausible or suspect, rather than as simply AI or not. The headline numbers in every report we produce use only the traffic we can verify or reasonably trust; the suspect traffic is shown separately, rather than counted in with the rest.
If a report you're given doesn't mention verification at all, it's worth asking the question, as it could mean the numbers are quite a way out from the reality.
HOW WE DO IT
4 Decide what counts
A few rules run through every query we make. They're the difference between measuring how much load the bots put on your server and measuring how much of your content they actually got, and none of them is complicated on its own. All of them are easy to miss.
- Requests aren't reads: A request is any line in the log. A read is a request where the server actually returned the page. A redirect (the server sent the visitor to a different address, usually because it changed) and a not-found are both requests, but neither is a read, because the visitor didn't get the content. Load figures count every request; reading figures count only the pages that were delivered, so a report that doesn't distinguish them is overstating.
- Today doesn't count: The current day is never complete and drags every average down, so we leave it out.
- Spam isn't AI: Anything probing for login pages or configuration files goes into a separate security view, never into the AI counts.
The 'reading figures' are what we mean when we talk about citability, whether you're the kind of source an AI system can and does reach for. (That's defined properly in our recent article Stop calling it "AI visibility".)
Once the counting rules are applied, we have numbers we can trust. The next question is what they can and can't tell you.
Not directly related, but worth knowing:
Before it does anything else on your site, every AI bot fetches a small file called robots.txt, which tells it which pages it's allowed to visit.
One thing we spotted when working on those counts was that robots.txt was redirecting on 90% of the sites we monitor, usually from the bare domain to the www version or from http to https. So every crawler was spending its first request on being sent to a different address, before it had read a single page.
Crawlers give each site a limited amount of attention per visit, so that was wasted attention on every visit from every bot, and it's a very quick thing to fix. So it's worth checking and updating if you're in the same position.
HOW WE DO IT
5 Know what the numbers can and can't say
A log line proves one thing: an AI system fetched the page. It does not prove that the page was cited, quoted, shown to anyone, or made any difference to the answer.
We describe what the logs show as reading, and we don't upgrade a read to a citation, because the logs can't support that and neither can anything else outside the operators' own systems.
Two things pull the live counts in opposite directions:
- Assistants remember: If five hundred people ask the same question in a week, ChatGPT may fetch your page once and reuse what it found for the other 499, so live fetches undercount how often your content is actually used.
- Not every fetch is discovery: If a person pastes a link to your page into ChatGPT and asks for a summary, that shows up in the log exactly the same as ChatGPT finding your page by itself.
We can't see the question or the person; we can only see the bot. So the most we can say concretely is that live fetches track AI reading in the right direction, not one for one.
Google is a separate scenario entirely:
When Google answers with an AI Overview or in AI Mode, it uses the copy of your site it already holds from ordinary search. There's no separate fetch in the logs and nothing to count. The only first-hand number for that is Google's own, in Search Console's generative AI report, which measures how often your pages appeared in Google's AI answers, not whether a bot visited.
It's a different dataset measuring a different thing, which is why we report it separately rather than mixing it into the log figures, and why we ask clients for Search Console access alongside the logs.
Be clear on how your 'Read to Referral' ratio is constructed:
And the figure people quote most, reads to referrals, isn't really a ratio. Across the sites we monitor it runs at roughly 140 reads for every visit an AI assistant sends through. But the reads come from the logs and the referrals come from GA4, which only counts people who consented and often files an AI-referred visit under "direct". They're two different sources measuring two different things. It's a useful signal for one site over time; it isn't a benchmark to compare across sites, and we don't.
We publish all of that in every report, in a section called "What this doesn't capture". It isn't small print; it's what lets the rest of the report go into a board pack without someone adding caveats later.
HOW WE DO IT
6 Read a month, interpret the results
Once the counts are clean, the view we use most is a journey. Each operator's bots move through your content in stages: train, index, read, represent, refer. Four of the five leave a trace. Here's one operator's month on the example site from our sample report:
| Stage | What it means | Where we see it | Month count |
|---|---|---|---|
| 1. Train | Copying your pages to train the model | Training crawler in the logs | 1,900 |
| 2. Index | Cataloguing your pages for the assistant to search | Search crawler in the logs | 70 |
| 3. Read | Fetching a page live to answer a question | Live fetches in the logs | 2,100 |
| 4. Represent | Naming or describing you inside the answer | Nowhere; only the operator can see this | No visibility |
| 5. Refer | A person clicking through from the answer | Referrals in the logs | 15 |
Stage 4 is the gap, and it's important to show it rather than leave it out, because the stages either side of it are both real.
What does the data mean?
What you're looking for is the gap between stages, because each gap points at a different kind of problem:
- Lots of training, few live reads: The operator has learned about you and isn't yet reaching for you. That's usually a plumbing question before it's a content one: is its crawler being blocked, and is your site in the index its assistant uses?
- Lots of live reads, hardly any referrals: The assistant is answering people's questions in the chat without sending them to you. That raises a choice about whether you want to give people a reason to click.
- Which pages get the live reads: This tells you what the assistants find useful on your site. On every site we monitor, it's the guides, explainers and research, not the service pages. That's a content brief in its own right, and a different article.
If you're going to try this yourself
Everything above can be done by a competent technical team, and I'd rather people looked at their logs than didn't. There are three things worth keeping in mind, as they will trip you up if you go in without expecting them.
The size:
A month of logs for a professional-services site is millions of lines, and the tools that store them aren't built to hand them over in bulk. You'll need somewhere to put them that can be queried at that scale, and you'll need to check, against the raw counts, that what you've collected covers everything; ours didn't for a while, and it wasn't easy to spot.
The fakes:
A meaningful share of what identifies itself as AI traffic isn't. Unless you're checking IP addresses against the operators' published lists, keeping those lists current, and being straight about the operators who don't publish one, your AI figures include traffic that probably isn't AI.
Knowing what you're looking at:
Requests aren't reads, redirects aren't visits, bulk crawlers aren't assistants, and Google's AI answers don't appear in the logs at all. Each of those is a small distinction that changes the number a lot.
If you'd rather get some help Take a look at our AI Citability Snapshot
We take one month of your server logs and show you which AI systems are reading your site, which pages they're reading, and how often they send someone back.
FAQs
Questions we get asked about reading AI crawler data
A live fetch is an AI assistant retrieving a specific page from your website at the moment a person asks it a question your content might answer. It shows up in your server logs as a request from a user-agent such as ChatGPT-User. It’s the closest observable signal to the moment of the question, and it’s different from a training crawl, which is bulk copying to train a model.
No. GA4 records a visit when a browser runs its tracking script and, in the UK, when the visitor consents. Bots do neither, so GA4 only ever sees the humans who clicked through from an AI answer and agreed to be counted. The server log records everything that touched the server, humans and bots, consented or not. The two are separate data sources and should be set side by side, never added together.
Not on its own. A user-agent is text that anyone can send, and ChatGPT-User is a common one to fake because many sites allow it through. The check is the IP address the request came from, compared with the addresses the operator publishes for its crawlers. Some operators publish these and some don’t, so some AI traffic can be verified and some can only ever be treated as plausible.
A request is any line in the server log, including redirects and missing-page responses. A read is a request where the server actually returned the page. Load figures count requests; reading figures count reads. A redirect is not a read.
No. A log line proves that an AI system fetched the page. It doesn’t prove the page was cited, quoted or shown to anyone. Only the operators can count appearances inside answers, and only Google publishes a figure for its own (the Search Console generative AI report). Logs describe reading, and we never describe a read as a citation.