Skip to content
AI Crawler Analytics

See which AI crawlers read your site — and prove which ones are real

Track every AI crawler that hits your pages — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more — verified against each vendor’s published IP ranges so a spoofed user-agent can’t fake its way in. Then close the gaps: map each blocked answer-engine to a concrete fix and apply it to your robots policy on WordPress in a click.

Credit card required · Cancel anytime · 7-day full access

Crawler Logs
18 crawlers · verifiedBusiness
Site health
7d28d90d
SEO Health
+6 this month
Search ConsoleClicks
1,284+18%
9.5K impressions
AI VisibilitySOV
38.0%+12pp
share of voice
Organic TrafficSessions
9,640+24%
this month
Feature health
Search Console · top moversView all
PageClicksPosStatus
/412#1.8Improving
/blog/ai-search-guide286#3.2Improving
/pricing204#2.4Healthy
/features/visibility171#4.1Healthy
Site Auditor
0Critical
12 warnings · audited 2h ago
BacklinksActive
128+14
Domain rating 41 · 30d
Content
28+6
articles · 30d
Keywords
345+27
saved in pool
Opportunities
1,049active
Quick 16Growth 518Strat 515
Traffic channels
  • Organic Search46%
  • Direct32%
  • Referral14%
  • AI answers8%

01The blind spot

Your analytics can’t see them. Your log can’t verify them.

Analytics runs on JavaScript, and crawlers don’t execute it — so the fastest-growing readers of your site never appear. Your server log does see them, but all it records is a user-agent string: a piece of text any script can type. The category has been counting text and calling it traffic.

  • Analytics tools measure people — AI crawlers never fire a pageview
  • A log line “from” GPTBot is a claim, not an identity — curl -A "GPTBot" produces one
  • Every unverified AI-crawler count you’ve seen is a count of strings, not visits
yoursite.com — access.log
What your server wrote down

203.0.113.7 - - [12/Aug/2026:03:41:07 +0000] "GET /blog/website-cost-pricing-guide HTTP/1.1" 200 48213 "-" "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)"

What anyone can type

$ curl -A "GPTBot" https://yoursite.com/

# that is all it takes to be “GPTBot” in an analytics tool

A name is not an identity

02Verified, not claimed

858 hits claimed GPTBot. 162 were telling the truth.

Last week, on a real account: 858 visits said GPTBot. SearchChamp checked each one against the source IP and OpenAI’s published ranges — 162 verified, 696 spoofed. Every one of the 78 hits claiming ClaudeBot was fake. A badge is only awarded when the address checks out, so the number you act on is the number that happened.

  • Verified · 162 — source IP confirmed to belong to the bot’s operator
  • Spoofed · 696 — the name said GPTBot; the source address didn’t
  • All 78 hits claiming ClaudeBot that week failed the check
GPTBotOpenAI · 858 visitsClaimed in your log
Verified · 162Spoofed · 696
162 FROM OPENAI’S PUBLISHED RANGES696 IMPOSTORS TYPING THE NAME
ClaudeBotAnthropic · 78 visitsClaimed in your log
Spoofed · 78
EVERY SINGLE HIT0 VERIFIED
Last 7 days · a real account · verified by source IP

03How the check works

Checked at the door. The raw IP is never written down.

Every AI-crawler hit is matched against the real source IP and the address ranges the crawler’s owner publishes — refreshed on a schedule, so a vendor adding IPs doesn’t silently turn real traffic into “spoof”. The check runs before the IP is hashed: verification happens at the door, and only the verdict is kept. When a crawler relies on DNS identity, a bounded reverse-DNS forward-confirm also finishes before storage.

  • Real source IP vs the operator’s published ranges — a name alone earns nothing
  • Ranges refresh on a schedule, so new vendor IPs don’t become false spoofs
  • The raw address is discarded once the verdict exists — checked, then hashed
1 · RequestGET /pricing
UA "GPTBot"
IP 203.0.113.7
as it arrives
2 · Range check203.0.113.7 ∈ OpenAI’s published ranges?refreshed on a schedule
3 · VerdictVerifiedbounded rDNS forward-confirm
4 · Storedverdict + sha256:a91f…c22eraw IP discarded — never stored
CHECK RUNS BEFORE THE IP IS HASHEDNEVER SLOWS AN INCOMING REQUEST
Verification at the door — only the verdict is kept

04Five outcomes, not two

Because “we don’t know” is three different things.

Most tools give you a binary: bot or not. SearchChamp reports five verdicts, because an unprovable identity isn’t the same as a fake one — and a check that hasn’t finished isn’t the same as a check that failed. When it can’t prove something, it tells you which kind of can’t-prove you’re looking at.

  • Unverifiable means the vendor publishes no ranges — not that the bot is fake
  • Pending check resolves; Check degraded is our outage on the record, not yours
  • A row with no recorded capture tier shows no badge at all — never a guess
VerifiedSource IP confirmed to belong to this bot’s operator.
SpoofedUser-agent claimed this bot but the source IP does not belong to its operator.
UnverifiableThis bot publishes no IP ranges, so its identity cannot be confirmed.
Pending checkVerification has not finished for this visit yet.
Check degradedThe verification service was unavailable, so this visit could not be checked.
UnknownNo verification status was reported for this visit.
It refuses to tell you something it can’t prove

05Three jobs, one company

Training, indexing, or fetching live — the job matters more than the name.

OpenAI alone runs GPTBot to gather training data, OAI-SearchBot to index pages for citation, and ChatGPT-User to fetch your page live when someone asks about you. Block the wrong one and you vanish from live answers while still feeding the training set. SearchChamp groups every crawler by what it actually does.

Breakdown by crawler

Last 7 days, grouped by what each bot does — and how many of its visits we could verify.

7 DAYS · 30 · 90

Training crawlers

Bots that fetch your content to train AI models.

GPTBot
OpenAI · 858 visits
Verified · 162Spoofed · 696
ClaudeBot
Anthropic · 78 visits
Spoofed · 78
Google-Extended
Google · 76 visits
Unverifiable · 76
meta-externalagent
Meta · 2412 visits
Unverifiable · 2412
Bytespider
ByteDance · 119 visits
Unverifiable · 119

Search-index crawlers

Bots that index your pages so AI assistants can cite them in answers.

Amazonbot
Amazon · 532 visits
Verified · 41Unverifiable · 487Pending check · 4
Applebot
Apple · 40 visits
Verified · 40
bingbot
Microsoft · 227 visits
Verified · 88Unverifiable · 116Pending check · 23
OAI-SearchBot
OpenAI · 186 visits
Verified · 101Spoofed · 85
Googlebot
Google · 181 visits
Verified · 167Spoofed · 14
PerplexityBot
Perplexity · 102 visits
Verified · 20Spoofed · 82

User-triggered fetchers

Bots that fetch a page live when a user asks an AI assistant about it.

ChatGPT-User
OpenAI · 422 visits
Verified · 198Spoofed · 224
Claude-User
Anthropic · 6 visits
Verified · 2Spoofed · 4
Blocking the wrong one is how sites disappear from live answers

06Every crawler, named

The full roster, by real registry token.

Every crawler is tracked by the token your log actually shows, vendor beside it — from GPTBot and ClaudeBot down to meta-externalfetcher and CCBot. And when something crawls you that isn’t in the registry, it lands in Unknown AI: recorded, never dropped. An unrecognised bot is a finding, not noise.

  • Real tokens like OAI-SearchBot (OpenAI) — the string your log actually shows
  • Sibling crawlers stay separate: Claude-User is not ClaudeBot
  • Unknown AI catches the rest — recorded, never dropped
OpenAI
GPTBotOAI-SearchBotChatGPT-User
Anthropic
ClaudeBotClaude-SearchBotClaude-User
Perplexity
PerplexityBotPerplexity-User
Google
GooglebotGoogle-Extended
Apple
ApplebotApplebot-Extended
Meta
meta-externalagentmeta-externalfetcher
Microsoft
bingbot
Amazon
Amazonbot
ByteDance
Bytespider
Common Crawl
CCBot
DuckDuckGo
DuckAssistBot
Mistral
MistralAI-User
Unknown AIAn unrecognised AI or generic bot user-agent is recorded, never dropped.
Named, not counted — and nothing falls off the list

07Capture, four ways

Something has to see the request. Pick what fits your stack.

Crawlers don’t run JavaScript, so capture happens where the request lands: edge middleware on Vercel or Next.js, a Cloudflare worker, the SearchChamp WordPress plugin, or a shipper that posts your existing access-log lines. The beacon is non-blocking and adds no page-speed cost; the log route changes nothing on your pages at all.

  • Ingest URL + secret token, sent in X-Ingest-Token — rotate it if it leaks
  • Non-blocking by design — detection never sits in the request path
  • Every row is stamped Beacon, Access log or Edge worker
Vercel / Next.jsCloudflareWordPressServer / access log
// middleware.ts — detects AI-bot user-agents, non-blocking
export function middleware(req) {
  reportAiBot(req)            // fire-and-forget beacon
  return NextResponse.next()  // page speed: untouched
}
INGEST URLhttps://ingest.searchchamp.com/v1/crawls
INGEST TOKENsc_live_••••••••••3f7aROTATE TOKEN

Sent in the X-Ingest-Token header. Keep it secret — rotate it if it leaks. Verify: request any page with a user-agent containing “GPTBot”.

EVERY VISIT STAMPED WITH ITS PIPELINE →BeaconAccess logEdge worker
Four ways in, one verified log

08An empty list that explains itself

“No bots yet” and “capture broken” are different facts.

A quiet dashboard is only reassuring if capture is provably alive. SearchChamp self-tests the tracker and tells you which state you’re in: working, not working, or not yet tested. An empty list under a passing self-test means no bots came — and it says so in those words.

  • A dated self-test proves the pipe is open before you trust a zero
  • “Too long since the last self-test” is called out, not papered over
  • Until the first test runs, the empty state says exactly what it can’t claim

Capture working

A self-test reached the tracker on Aug 12, 2026. Bot visits will be recorded when they arrive.

Capture not working

The last self-test was on Jun 3, 2026 — too long ago. New bot visits may not be recorded. Re-check your tracking snippet.

THE LIVE ACCOUNT, TODAY

Capture not yet tested

No self-test has run yet. Once your snippet is live, an empty list will mean no bots have visited — not that capture is broken.

Three states — a zero you can trust

09Windows, honestly

The window you see is the window you actually have.

Every view runs on a 7 / 30 / 90-day toggle, with your plan’s real retention beside it: Starter keeps 7 days of crawl history, Pro 90, Agency the full record. A locked window says why it’s locked. And if your plan can’t be determined, SearchChamp falls back to the shortest window — it never flatters you with data it shouldn’t show.

  • Locked windows read “Available on higher plans” — no silent empty charts
  • Unknown plan? Fail closed: the most restrictive window, never the longest
  • Past the daily event cap, a visible sampling notice — not silent truncation
7 days30 daysLOCKED90 daysLOCKED

Available on higher plans — your Starter plan includes a shorter history window.

STARTER7 days of crawl historyYOUR PLAN
PRO90 days of crawl history/pricing
AGENCYFull history/pricing
Sampling active — this site has passed its daily limit of 50,000 events; counts above the cap are sampled.

UNKNOWN PLAN → FALLS BACK TO THE MOST RESTRICTIVE WINDOW. A LONGER ONE IS NEVER LEAKED BY ACCIDENT.

Retention, stated — not implied

10Crawled vs cited

Your three most-crawled pages can never be cited.

The reconciliation most tools never run: pages verified AI crawlers fetch, against pages AI engines actually cite. On this account the top three crawled targets are an image endpoint, robots.txt and the sitemap — 2,839 fetches of things no answer will ever quote. Meanwhile one page is being cited without a recent verified crawl at all.

Crawled vs. cited

How the pages AI crawlers fetch line up with the pages AI engines actually cite.

CRAWLED BUT NEVER CITEDCRAWLS
THE TOP THREE — AN IMAGE ENDPOINT, ROBOTS.TXT, THE SITEMAP — ARE CRAWL BUDGET THAT CAN NEVER BECOME A CITATION.
/_next/image1447Crawled by ChatGPT-User, ClaudeBot, GPTBot, OAI-SearchBot
/robots.txt852Crawled by ClaudeBot, Claude-User, OAI-SearchBot, PerplexityBot
/sitemap.xml540Crawled by ClaudeBot, GPTBot
/services/ai-solutions/ai-automation-workflows45
/services/ai-solutions/ai-chatbot-development41
/blog/ai-chatbot-development-cost31
/services/ecommerce/shopify-development31
/services/mobile-development30
/blog/website-cost-pricing-guide27
CITED BUT NEVER CRAWLEDCITATIONS
/solutions/ecommerce/shopify-development1Cited by perplexity

Quoted without a recent verified crawl — the engine is answering from what it already has. Check the page stays reachable and fresh.

Reconciles pages verified AI crawlers fetched against pages AI engines cited. “Crawled, never cited” pages are read but not quoted — improving extractability raises (not guarantees) citation odds. “Cited, never crawled” pages are quoted without a recent verified crawl — check they stay reachable and fresh.

11A floor, not a ceiling

Crawl volume is loud. Attributable sessions are a floor.

Verified AI-bot crawl volume against real AI-referred sessions — with the caveat most tools quietly omit: many AI apps strip the referrer, so GA4 files those visits under “(direct)”, and Bing Copilot can’t be told apart from Bing organic at all. SearchChamp filters to genuinely attributable AI sources and labels the count for what it is.

  • GA4 filtered to AI-assistant sources — not the whole “Referral” channel
  • Measured = reliable UTM/referrer · Referrer-only = bare host, likely undercounted
  • The undercount is printed on the panel, not buried in the docs
VERIFIED CRAWLS · 7D819IP-verified AI-bot visits
VS
AI-SOURCE SESSIONS · 7D63GA4, AI-assistant sources onlyA FLOOR, NOT A CEILING
AI-assistant referral traffic is undercounted here: many AI apps (native mobile clients, in-app browsers) omit or strip the referrer, so GA4 reports those visits as “(direct)” instead of an attributable AI source. Bing Copilot referrals cannot be attributed from GA4’s session source. Treat this count as a floor, not a ceiling.

Source: GA4 (AI-source filtered) + verified crawler logs · As of Aug 14, 2026

ChatGPTMEASURED
Microsoft CopilotREFERRER-ONLY
Google AI ModeREFERRER-ONLY
GeminiREFERRER-ONLY

MEASURED — reliably carries an attributable UTM/referrer · REFERRER-ONLY — GA4 only sees a bare referring host; likely undercounted

Attribution quality, stamped per source

12From crawl to conversion

Crawled → Cited → Visible → Clicked → Converted.

The full connected funnel — and the three honest states a stage can be in. No source connected: a specific connect CTA, never a fabricated 0. Source connected but the live pull failed: “Temporarily unavailable” — a different thing. And a low-sample stage carries its confidence interval, so a ratio inside the error bars reads “Within noise” instead of a confident arrow.

  • Real 30-day values: Crawled 3,719 · Cited 7 · Visible 2.0%
  • “Connect GA4 to see clicks” is a state, not an error — and never a fake zero
  • No conversion rate across Visible: it’s a share-of-voice %, not a count

From crawl to conversion

How visibility with AI crawlers turns into citations, clicks, and revenue.

7d30d90d
Crawled3,719Verified AI-bot visits
Verified AI-bot visits, last 30 days.
0%
Cited7Prompts with a brand citation
Prompts with a brand citation in the selected window.
WITHIN NOISE
Visible2.0%n=277 · ±2.6pp
Share of voice across your tracked prompts.
ClickedTEMPORARILY UNAVAILABLEGA4 is connected — the pull failed
AI-source GA4 sessions, last 30 days.
ConvertedTEMPORARILY UNAVAILABLEKey events + Shopify orders
GA4 key events (+ Shopify orders, where connected).
A NUMBER — the stage has verified data in the selected window.
A CONNECT CTA — “Connect GA4 to see clicks.” No source connected; never a fabricated 0.
TEMPORARILY UNAVAILABLE — source connected, live pull failed. Not a CTA, not a zero.
Ratio pips read 0%, — , or “within noise” — never a confident arrow on a small sample

13Can they even reach you?

Everything above is measured. This is probed — and labelled.

The verified log shows what really happened. The Access Guard answers the forward-looking question — could each answer engine reach you if it tried? — by probing your site with each bot’s user-agent. A probe is a forecast, not a verdict, so every single row carries a heuristic tag and a confidence level, and the evidence says exactly how it was measured.

100
AI ACCESS SCORE

AI Access Score 100/100. No AI crawlers appear to be blocked. An llms.txt file is present. Probe results are heuristic: measured with a bot user-agent from SearchChamp infrastructure, not the crawler’s verified IP.

LLMS.TXT PRESENTSCANNED AUG 14, 2026
↻ RESCAN

AI answer & search

Crawlers that power AI answers and cite sources (ChatGPT, Claude, Perplexity, …)

ChatGPT-UserHEURISTICHIGH CONFIDENCEALLOWED
Claude-SearchBotHEURISTICMEDIUM CONFIDENCEALLOWED
Claude-UserHEURISTICMEDIUM CONFIDENCEALLOWED
DuckAssistBotHEURISTICLOW CONFIDENCEALLOWED
MistralAI-UserHEURISTICLOW CONFIDENCEALLOWED
OAI-SearchBotHEURISTICHIGH CONFIDENCEALLOWED
Perplexity-UserHEURISTICMEDIUM CONFIDENCEALLOWED
PerplexityBotHEURISTICHIGH CONFIDENCEALLOWEDEVIDENCE ▾

EVIDENCE — PERPLEXITYBOT

bot user-agent served 200 with content. Heuristic: measured from SearchChamp infrastructure with a bot user-agent only. Real crawlers are verified by IP/DNS/Web-Bot-Auth against the vendor’s published ranges, so this reflects how your site responds to that user-agent, not a guaranteed verdict for the real bot.

Classic search

Traditional web-search crawlers (Google, Bing)

bingbotHEURISTICHIGH CONFIDENCEALLOWED
GooglebotHEURISTICHIGH CONFIDENCEALLOWED

Model training

Crawlers that fetch content to train AI models.

AmazonbotHEURISTICMEDIUM CONFIDENCEALLOWED
BytespiderHEURISTICLOW CONFIDENCEALLOWED
CCBotHEURISTICMEDIUM CONFIDENCEALLOWED
Verified-by-IP and probed-by-user-agent are different things — the page says which is which on every row

14Fix it, then re-check

Every fix shows its diff. Every “done” gets re-scanned.

Findings become fixes you can read before they happen: the proposed robots.txt directives and the exact resulting file. Apply to WordPress in one click, get steps for what can’t be automated — Cloudflare rules included — and marking a fix done queues a re-scan to verify it took. All under a standing reminder: robots.txt is advisory, and “Not addressed” was never an allow.

  • Three directive states: Disallowed · Allowed · Not addressed — never flattened to two
  • “Applied — pushed to your WordPress site.” Or honestly: no live connection to push to yet
  • Publishing a preference is stated as a preference — honoured by well-behaved bots, not enforced

Fix access issues

Turn findings into fixes — apply what can be automated, get steps for the rest.

FIND FIXES

Publish an explicit allow for answer-engine crawlers

PROPOSED ROBOTS.TXT DIRECTIVES → RESULTING ROBOTS.TXT
  User-agent: *
  Allow: /
  Sitemap: https://yoursite.com/sitemap.xml
+ User-agent: OAI-SearchBot
+ Allow: /
+ User-agent: PerplexityBot
+ Allow: /
APPLY FIXGET INSTRUCTIONSMARK DONE
Applied — pushed to your WordPress site.
Marked done — a re-scan has been queued to verify.
DisallowedAsks this bot to stay out. Well-behaved crawlers honour it — it is not enforced.
Allowedrobots.txt explicitly addresses this bot and does not block it.
Not addressedNever mentioned — no decision has been made. This is not an allow.
ROBOTS.TXT IS ADVISORY, NOT ENFORCEMENT
Read the diff, apply, re-scan — nothing is taken on faith
Common queries

AI Crawler Analytics, answered.

Hover or click a question for the answer
01

Every AI crawler hit on your site — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more — classified as verified, spoof, unverifiable, or pending, so you can see which answer engines are really reading your content instead of guessing from a user-agent string.

GENERAL
02

A bot operated by an AI company that fetches web pages so its models and answer engines can use them — either to train on, or to retrieve a page while answering someone’s question right now. It is the AI-era equivalent of Googlebot: if it cannot reach and read your page, that engine has nothing of yours to cite.

GENERAL
03

Every crawler hit is matched against the real request IP and checked against the crawler owner’s published address ranges. A bot name alone never earns a verified badge — a spoofed user-agent that does not come from the real IP range shows as spoof, not verified. On a real account in one recent week, 696 of 858 hits claiming to be GPTBot were spoofed.

CAPABILITIES
04

GPTBot crawls the web broadly to gather content for training. ChatGPT-User is the fetch that happens when someone asks ChatGPT a question and it retrieves your page to answer it right then. They are separate user-agents you can allow or block independently — and the distinction matters, because blocking the second is what removes you from live answers, while blocking the first only affects training. SearchChamp reports them separately rather than collapsing both into “OpenAI”.

CAPABILITIES
05

Something has to see the request. Choose whichever fits: a few lines of edge middleware on Vercel or Next.js, a Cloudflare worker, the SearchChamp WordPress plugin, or a shipper that posts your existing server access-log lines. The beacon is non-blocking and adds no page-speed cost, and the access-log route changes nothing on your pages at all.

INTEGRATION & SCALE
06

Yes — each blocked answer engine is mapped to a concrete fix. You can see the proposed robots.txt directives and the resulting file before anything happens, apply it to WordPress in one click, and a re-scan is queued to confirm it took effect. For anything that cannot be automated, such as a Cloudflare rule, you get the steps. Worth knowing: robots.txt is advisory, not enforcement — well-behaved crawlers honour it, and nothing forces them to.

INTEGRATION & SCALE
07

Verified crawler tracking is part of the core platform on every paid plan, not a paid add-on. What differs is how much crawl history is retained: Starter keeps 7 days, Pro 90 days, and Agency the full record. See the pricing page for what each plan costs.

PRICING
08

No. Crawler verification runs continuously as part of monitoring a connected site rather than being billed per check, so you are not choosing between watching closely and controlling spend.

PRICING

Know exactly which AI engines read — and cite — your site

7-day free trial · credit card required · cancel anytime