Pillar guide

The Authority Engine: How We Build Sites AI Engines Cite

The short answer

How do you build a site AI engines trust?

An authority engine is a publishing system, not a pile of posts: demand research that is allowed to say no, a content contract enforced by build gates, production at scale — 400+ fact-audited pages across our 3 production brands [our data] — URL submission through IndexNow, and measurement through GA4 AI-referral channels and Search Console mining. Trust is earned passage by passage, and no step in the system promises a citation.

Most of what is written about AI search is commentary — descriptions of the phenomenon by people who do not operate sites through it. This page documents the alternative: the operating system we run across 3 production brands and 400+ published, fact-audited pages [our data]. It is the system that produced this library, described plainly enough to copy.

One sentence of honesty before the machinery: no step below controls what any answer engine does. The system's bet is narrower — that pages engineered to be checkable and liftable earn citations more often than pages that are not — and every claim it makes about results comes from our own logs, not from projection.

What is an authority engine?

An authority engine is the whole pipeline that turns a niche into a maintained library: demand research, site architecture, contracted content production, a fact-audit gate, indexing plumbing, and measurement. The output is not "content" but an asset with known properties — every page answer-first, every claim sourced, every review date enforced.

StageOutputThe gate
Demand researchQuestion inventory, competitor censusAllowed to conclude "don't build"
ArchitectureClusters: pillar, learn, glossary, scenarios, costEvery page has a job
ProductionAnswer-first MDX under the content contractContract violations fail the build
Fact auditClaim-by-claim verificationUnsourced claims do not ship
IndexingSitemaps, IndexNow submission on publishMachine-checkable surfaces
MeasurementGA4 AI channel, GSC mining, prompt checksResults claimed only from data

The stages are ordinary. The discipline of running all 6, in order, with mechanical enforcement, is what the census of this niche shows almost nobody doing.

How does demand research decide what gets built?

The research pass inventories the questions a niche actually asks — ours for this site logged 92 distinct phrasings across definitional, how-to, cost, and skeptical intents [our data] — then censuses who currently answers them and how well. The output is an architecture where every planned URL maps to a real question, and an honest read on whether the niche is winnable at all.

The feature that makes the stage worth paying for is that it can return "no." A niche whose head terms are owned by platforms with data we cannot beat, and whose long tail lacks volume, fails the pass — and the correct output is to not build. Most of the money wasted on content is spent learning that fact the expensive way.

How does the content contract keep pages citable?

The contract is a written spec that makes citability mechanical instead of aspirational. Its load-bearing rules: a 40–75-word direct answer at the top of every page; question-shaped headings whose sections stand alone; at least 1 real table; and a bright line on evidence — every statistic either cites a URL from a closed list of primary sources or is marked [our data] and named to one of our builds. Nothing else ships. The compliance floor is absolute: no page promises rankings, citations, or traffic, because nobody controls answer-engine output.

Enforcement is the difference between a style guide and a system. Our build gates parse every page before publish: a direct answer outside its word bounds, a missing source, a stale review date, or a promise-shaped sentence fails the build. When we ran a site-wide audit under this contract, it checked 445 claims and corrected 40 [our data] — a 9% correction rate on our own shipped content, and the reason the audit stage is not optional. The per-page economics of that discipline are broken down in what a fact-audited page costs.

The page-level techniques the contract encodes — extractable passages, answer blocks, the evidence tiers — are the subject of the operator's guide to GEO, and the academic grounding for structure-plus-citations comes from the Princeton GEO benchmark (arXiv 2311.09735, KDD 2024).

How do pages get discovered and indexed?

Three surfaces, all machine-first. Everything ships as static, crawlable HTML — most AI crawlers execute no JavaScript, so a page rendered client-side does not exist to them. Google's guidance for its AI features points at the same fundamentals: its AI experiences are built on the same crawling, indexing, and helpfulness systems as Search (Google, May 2025; ai-optimization-guide), which is why the engine optimizes each page once for both surfaces.

Second, submission: publishes fire an IndexNow request so participating engines learn about changed URLs immediately — up to 10,000 per request (indexnow.org) — with the full procedure in the IndexNow setup guide. Third, machine surfaces: XML sitemaps segmented by cluster, plus an llms.txt file per the llms.txt proposal (llmstxt.org) — not an adopted standard, and no major engine has committed to reading it, which is exactly why we deployed it as an experiment with logging rather than as a faith gesture.

How do we measure whether any of it works?

With infrastructure built before claims are made, all shipped on our own brands first [our data]: a GA4 channel group that isolates AI-assistant referrals (the build is documented in tracking ChatGPT traffic in GA4), a Search Console miner that approximates AI Overview exposure from query and impression patterns, and manual prompt checks run across the major engines.

Measurement is also where the receipts come from. Our documented Google AI Overview citation — the query class, the page's shape, and how the win was observed — is published as a full scenario, because a claim with its evidence attached is the only kind this system is allowed to make.

What does the system refuse to do?

It refuses to promise outcomes — no citation, ranking, or traffic result is guaranteed by any step above, and a vendor telling you otherwise is describing a system they do not control. It refuses to launder statistics: if a number is not on our closed source list or in our logs, it does not ship, however widely it circulates. And it refuses to hide failures — the experiment that moves nothing gets published with its logs, because negative results are the one content type this niche cannot fake.

Refused fastest of all: building where the research says not to. Most niches we examine do not clear the bar, and the system's cheapest output is the build that never happens.

Frequently asked questions

How do you build a site AI engines trust?

Run a system, not a posting schedule: research the demand first, write under a content contract that forces every claim to a source or to first-party data, gate publication mechanically, submit URLs on publish, and measure with your own analytics. Our version runs 400+ audited pages across 3 brands [our data].

What is a content contract?

A written spec every page must pass before it ships: answer-first structure, a 40–75-word direct answer, question-shaped headings, a closed list of citable sources, and compliance rules like never promising rankings. Ours is enforced by build gates — a violating page fails the build.

How many pages does an authority site need?

Fewer than volume sellers claim, more than a blog produces by accident. Our production builds each shipped their first 100–300 content files inside their first month [our data] — but every page must survive fact-checking, because scale multiplies whatever quality you actually have.

How long before an authority build shows results?

Treat 90 days as the earliest honest read of trend, and expect citation appearances before referral traffic. On our own fleet we documented a Google AI Overview citation through dated manual checks [our data] — observed within days of the page shipping, but measured, not promised.

Sources

  1. Top ways to ensure your content performs well in Google's AI experiencesGoogle
  2. Google's guide to optimizing for generative AI featuresGoogle
  3. GEO: Generative Engine Optimization (Princeton et al., KDD 2024)arXiv
  4. IndexNow protocol documentationIndexNow
  5. The llms.txt proposalllmstxt.org