Pillar guide
The AI Visibility Audit: How to Check What Answer Engines Say About You
The short answer
How do I audit my brand's visibility in AI answers?
An AI visibility audit is a repeatable evidence pass, not a score. This one has 5 passes: a frozen brand prompt battery sampled across engines, a verified crawler-log review, a citation inventory built from dated manual SERP checks, a trust-liability sweep, and a fix-priority pass. We publish it as protocol rather than as a track record — our builds are weeks old — and none of it promises a citation.
Most brands asking "what does ChatGPT say about us?" want a number. The honest deliverable is different: a dated evidence file, five passes deep, that shows what answer engines currently say, whether they can even reach your pages, what they cite instead, and which liabilities on your own site are working against you.
This page documents that method, assembled from the passes we run on our own 3 production builds [our data]. One disclosure before the procedure: our fleet is young — three builds, not a decade of engagements — so what follows is the method and its limits, not a track record, and several passes below have run on one build rather than all of them. Nothing here produces citations, rankings, or traffic, and any audit sold on that promise is selling a system nobody operates.
What is an AI visibility audit?
An AI visibility audit is a structured evidence pass over four separate questions: what engines say about you, whether their crawlers can reach you, what gets cited on your topics, and what on your site is actively untrustworthy. It ends in a fix order. It does not end in a score, because no calibrated score exists to report.
| Pass | Question it answers | Evidence it produces | Its hard limit |
|---|---|---|---|
| 1. Prompt battery | What do engines say about us? | Cited / mentioned / absent / wrong, per engine, per date | A sample of a non-deterministic system |
| 2. Crawler logs | Can engines fetch us at all? | Verified bot fetches by user agent and path | Fetching is not citing |
| 3. Citation inventory | Who gets cited on our questions? | Dated SERP and answer screenshots, GSC query patterns | No tool has an AI Overview dimension |
| 4. Trust sweep | What on our site is a liability? | Defect ledger: markup, claims, contradictions | Finds defects, not their impact |
| 5. Fix order | What do we do first? | Ranked backlog with owners | Priority is judgment, not measurement |
The passes run in that order for a reason. Passes 1 and 3 tell you where you stand; pass 2 tells you whether the problem is even about content; pass 4 finds the damage that makes every later effort worth less. Skipping to content work before pass 2 and pass 4 is the most common way an audit wastes money.
What does the brand prompt battery test?
The battery tests what engines actually say when someone asks about your category, your competitors, and you — measured as a frozen list of 20 to 30 prompts, run in each engine, on a noted date, and scored four ways: cited, mentioned, absent, or wrong. "Wrong" is the column that earns the whole exercise, because a confidently misstated fact about your business is a liability with a name and a date attached.
Freeze the prompt list before the first run. A battery that drifts month to month measures your prompt writing rather than your visibility. The build method, the four prompt bands, and the scoring sheet are documented in full in how to measure AI share of voice for free; this pass is that method used as an audit input.
Two limits belong in the report next to the numbers. First, engines are non-deterministic — the same prompt can return different answers on the same day, so a single flipped result is noise and three consecutive months of the same flip is signal. Second, some answers never touch your site at all: Kevin Indig's 2026 analysis found 24% of ChatGPT answers were generated without fetching any live page (Growth Memo, 2026). Part of what an assistant says about you today is frozen in training data, and nothing you publish this quarter reaches it quickly.
What do server logs prove that analytics cannot?
Logs prove reachability, which is the only part of this audit with hard ground truth. A verified crawler either requested your URL or it did not — no sampling, no interpretation. Analytics cannot show you this at all: the documented AI crawlers fetch raw HTML and do not execute JavaScript, so they never run an analytics script.
Pull the log lines by user agent, verify each one against the platform's published verification method, and count fetches by bot and by path. The full procedure, including the reverse-DNS and IP-range checks that separate real bots from impersonators, is in AI bot analysis from server logs. What the counts tell you is narrow and useful: which of your sections are being read, which are invisible, and whether a robots rule you forgot is doing the deciding.
What the counts do not tell you is value. Cloudflare's crawl-to-click analysis measured wildly different extraction economics per platform in July 2025 — roughly 5.4 pages fetched per referral click for Google, ~1,091:1 for OpenAI, ~195:1 for Perplexity and ~38,066:1 for Anthropic (Cloudflare, 2025). Heavy crawling is not a compliment, and light crawling is not a verdict. Record the ratio, do not celebrate it.
How do you inventory citations when no tool has the dimension?
By hand, with dates — because Search Console has no AI Overview dimension and no third-party tool has access to one either. The inventory is a dated record of which sources an engine cited for each battery prompt, screenshotted, plus the Search Console query patterns underneath the pages you care about.
Google documents eligibility mechanically: a page must be indexed and eligible to appear as a snippet, with no snippet controls blocking extraction (Google, AI Features). Check that first — it is a five-minute pass that occasionally explains everything. Then mine the query side: on our insurance lead-gen build, Search Console returned actionable query data by day 4, and on our auto-finance rebuild by day 6 because the domain already had history [our data].
This is also the pass that recorded the one AI Overview citation our fleet can document: a glossary page observed as a cited source through dated manual checks of the live result, days after it shipped [our data]. We record that as an observation with an n of 1, not as a benchmark, and we say so every time it appears on this site.
Set expectations about what a citation is worth while you are here. Pew's browsing study found users clicked a cited source in about 1% of visits that included an AI summary, with overall click rates of 8% on pages with a summary against 15% without (Pew Research Center, 2025). An audit that treats citations as a traffic forecast is misreading the channel.
What trust liabilities have to be fixed before anything else?
The trust sweep looks for things on your own site that make it less citable and more penalizable, and it takes priority over every content idea in the backlog. The checklist that matters: markup that describes things that do not exist, text hidden from users but served to crawlers, the same fact stated differently on different pages, and consent or policy copy that contradicts what the site actually does. Google's Search Essentials covers the spam-policy half of that list directly.
We learned to run this first from our auto-finance rebuild, which arrived with inherited liabilities: fabricated review markup, hidden-text blocks, 4 contradictory approval-rate claims live at the same time, and a consent defect. All of it was removed or replaced — one approval-rate claim with a published methodology page, sitemap junk purged — before the apex ever pointed at new content [our data]. The rule that came out of it: liabilities poison an authority build, so trust repair is a prerequisite, not an option.
Contradictions deserve their own line, because they are the failure mode that grows with your library. Where your own pages answer one question three different ways, an engine has no way to choose the right one and every correction you publish competes with your own archive. The workflow for fixing what engines already say wrong about you is in how to fix what AI says about your brand.
How do you decide what to fix first?
By cost of delay, not by opportunity size. The order we use is mechanical before editorial: reachability, then liabilities, then eligibility, then structure, then coverage.
| Priority | Fix class | Why it outranks the next one |
|---|---|---|
| 1 | Crawler reachability, blocked paths | Nothing downstream matters if the fetch never happens |
| 2 | Trust liabilities and contradictions | They devalue every page you publish afterward |
| 3 | Snippet eligibility and indexing | Documented prerequisite for AI feature appearance |
| 4 | Passage structure on existing pages | Cheapest per-page work with a plausible mechanism |
| 5 | New coverage on uncovered questions | Highest effort, slowest feedback |
Most audit reports invert this, because new content is the most sellable finding. Against our own interest as a shop that builds content libraries: if your audit's first four rows are non-empty, a new content program is the wrong purchase this quarter. The operating system that all five passes feed into — research, contract, gates, indexing, measurement — is documented in the Authority Engine.
What can this audit not tell you?
It cannot tell you why. Answer engines do not publish selection reasons, and no outside observer can isolate the variable that produced a citation or its absence. An audit narrows the field by ruling causes out — the crawler was blocked, the page was not snippet-eligible, the claim contradicted three other pages — and that is a genuinely useful thing to be able to say.
It also cannot tell you what happens next. Every pass here is a sample or a snapshot, the systems change without notice, and none of the work described on this page is a promise of citations, rankings, or traffic. If you would like us to run it on your site and say plainly whether a build is even warranted, that question is what the fit check exists to answer — including the answer where the honest recommendation is not to build.
Frequently asked questions
How do I audit my brand's visibility in AI answers?
Run 5 passes: a frozen prompt battery across engines, a verified crawler-log review, a citation inventory from dated manual SERP checks and Search Console query mining, a trust-liability sweep of your existing pages, and a fix-priority pass. Each pass answers a different question.
What tools do I need for an AI visibility audit?
A spreadsheet, access to your server or CDN logs, and Search Console. Paid AI-visibility trackers sample the same non-deterministic systems more often; they do not see ground truth. Start manual, and buy frequency later only if the battery outgrows an afternoon.
Can an audit tell me why I am not being cited?
It can rule causes out, not in. Logs prove whether crawlers reached you; snippet controls prove whether you are eligible; the battery shows what engines say instead. What no audit can isolate is why a specific engine chose a specific source on a specific day.
How often should the audit run?
The prompt battery and citation inventory suit a monthly cadence, because answer engines are non-deterministic and only trends mean anything. Crawler logs are worth a look whenever traffic or robots rules change. The trust sweep is a one-time prerequisite you repeat after inheriting a domain.
Does an audit fix my AI visibility?
No. An audit produces evidence and a fix order; publishing and repair produce change, and no step controls what an answer engine outputs. Anyone selling an audit that ends with a promised citation count is selling a result they cannot deliver.
What is the first thing an audit usually finds?
Something mechanical. Blocked or unverified crawlers, pages that are not snippet-eligible, and contradictory claims are what passes 1 through 4 are ordered to surface first — all cheaper to fix than the content work everyone wants to start with.
Sources
- AI Features and Your Website — Google
- Google Search Essentials — Google
- The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals — Cloudflare
- Google users are less likely to click on links when an AI summary appears — Pew Research Center
- State of AI Search Optimization 2026 — Growth Memo (Kevin Indig)