Pillar guide
The Authority Engine: How We Build Sites AI Engines Cite
The short answer
How do you build a site AI engines trust?
An authority engine is a publishing system, not a pile of posts: demand research that is allowed to say no, a content contract enforced by build gates, production at scale — 400+ fact-audited pages across our 3 production brands [our data] — URL submission through IndexNow, and measurement through GA4 AI-referral channels and Search Console mining. Trust is earned passage by passage, and no step in the system promises a citation.
Most of what is written about AI search is commentary — descriptions of the phenomenon by people who do not operate sites through it. This page documents the alternative: the operating system we run across 3 production brands and 400+ published, fact-audited pages [our data]. It is the system that produced this library, described plainly enough to copy.
One sentence of honesty before the machinery: no step below controls what any answer engine does. The system's bet is narrower — that pages engineered to be checkable and liftable earn citations more often than pages that are not — and every claim it makes about results comes from our own logs, not from projection.
What is an authority engine?
An authority engine is the whole pipeline that turns a niche into a maintained library: demand research, site architecture, contracted content production, a fact-audit gate, indexing plumbing, and measurement. The output is not "content" but an asset with known properties — every page answer-first, every claim sourced, every review date enforced.
| Stage | Output | The gate |
|---|---|---|
| Demand research | Question inventory, competitor census | Allowed to conclude "don't build" |
| Architecture | Clusters: pillar, learn, glossary, scenarios, cost | Every page has a job |
| Production | Answer-first MDX under the content contract | Contract violations fail the build |
| Fact audit | Claim-by-claim verification | Unsourced claims do not ship |
| Indexing | Sitemaps, IndexNow submission on publish | Machine-checkable surfaces |
| Measurement | GA4 AI channel, GSC mining, prompt checks | Results claimed only from data |
The stages are ordinary. The discipline of running all 6, in order, with mechanical enforcement, is what the census of this niche shows almost nobody doing.
How does demand research decide what gets built?
The research pass inventories the questions a niche actually asks — ours for this site logged 92 distinct phrasings across definitional, how-to, cost, and skeptical intents [our data] — then censuses who currently answers them and how well. The output is an architecture where every planned URL maps to a real question, and an honest read on whether the niche is winnable at all.
The feature that makes the stage worth paying for is that it can return "no." A niche whose head terms are owned by platforms with data we cannot beat, and whose long tail lacks volume, fails the pass — and the correct output is to not build. Most of the money wasted on content is spent learning that fact the expensive way.
How does the content contract keep pages citable?
The contract is a written spec that makes citability mechanical instead of aspirational. Its load-bearing rules: a 40–75-word direct answer at the top of every page; question-shaped headings whose sections stand alone; at least 1 real table; and a bright line on evidence — every statistic either cites a URL from a closed list of primary sources or is marked [our data] and named to one of our builds. Nothing else ships. The compliance floor is absolute: no page promises rankings, citations, or traffic, because nobody controls answer-engine output.
Enforcement is the difference between a style guide and a system. Our build gates parse every page before publish: a direct answer outside its word bounds, a missing source, a stale review date, or a promise-shaped sentence fails the build. When we ran a site-wide audit under this contract, it checked 445 claims and corrected 40 [our data] — a 9% correction rate on our own shipped content, and the reason the audit stage is not optional. The per-page economics of that discipline are broken down in what a fact-audited page costs.
The page-level techniques the contract encodes — extractable passages, answer blocks, the evidence tiers — are the subject of the operator's guide to GEO, and the academic grounding for structure-plus-citations comes from the Princeton GEO benchmark (arXiv 2311.09735, KDD 2024).
How do pages get discovered and indexed?
Three surfaces, all machine-first. Everything ships as static, crawlable HTML — most AI crawlers execute no JavaScript, so a page rendered client-side does not exist to them. Google's guidance for its AI features points at the same fundamentals: its AI experiences are built on the same crawling, indexing, and helpfulness systems as Search (Google, May 2025; ai-optimization-guide), which is why the engine optimizes each page once for both surfaces.
Second, submission: publishes fire an IndexNow request so participating engines learn about changed URLs immediately — up to 10,000 per request (indexnow.org) — with the full procedure in the IndexNow setup guide. Third, machine surfaces: XML sitemaps segmented by cluster, plus an llms.txt file per the llms.txt proposal (llmstxt.org) — not an adopted standard, and no major engine has committed to reading it, which is exactly why we deployed it as an experiment with logging rather than as a faith gesture.
How do we measure whether any of it works?
With infrastructure built before claims are made, all shipped on our own brands first [our data]: a GA4 channel group that isolates AI-assistant referrals (the build is documented in tracking ChatGPT traffic in GA4), a Search Console miner that approximates AI Overview exposure from query and impression patterns, and manual prompt checks run across the major engines.
Measurement is also where the receipts come from. Our documented Google AI Overview citation — the query class, the page's shape, and how the win was observed — is published as a full scenario, because a claim with its evidence attached is the only kind this system is allowed to make.
What does the system refuse to do?
It refuses to promise outcomes — no citation, ranking, or traffic result is guaranteed by any step above, and a vendor telling you otherwise is describing a system they do not control. It refuses to launder statistics: if a number is not on our closed source list or in our logs, it does not ship, however widely it circulates. And it refuses to hide failures — the experiment that moves nothing gets published with its logs, because negative results are the one content type this niche cannot fake.
Refused fastest of all: building where the research says not to. Most niches we examine do not clear the bar, and the system's cheapest output is the build that never happens.
Frequently asked questions
How do you build a site AI engines trust?
Run a system, not a posting schedule: research the demand first, write under a content contract that forces every claim to a source or to first-party data, gate publication mechanically, submit URLs on publish, and measure with your own analytics. Our version runs 400+ audited pages across 3 brands [our data].
What is a content contract?
A written spec every page must pass before it ships: answer-first structure, a 40–75-word direct answer, question-shaped headings, a closed list of citable sources, and compliance rules like never promising rankings. Ours is enforced by build gates — a violating page fails the build.
Do you need different content for AI search and Google search?
No. Google's May 2025 guidance says its AI experiences build on the same crawling, indexing, and helpfulness systems as Search, and the extractable structure answer engines lift is the same structure that wins snippets. The engine produces 1 page for both surfaces.