Guide
Should You Publish Experiments That Didn't Work?
The short answer
Should you publish experiments that didn't work?
Yes, when you can state the method and its limits. We publish 2: speakable schema ships on every page of 3 production builds with no effect we can attribute, and our llms.txt watch has logged 0 requests from documented AI engine crawlers so far, in a window still running. A null result on a small sample is never proof of absence, and the page has to say so.
The experiment that moved nothing is the most useful page nobody publishes. Marketing content selects hard for successes, which is exactly why a documented failure carries information: it is the one content type this niche cannot fake cheaply, because faking a null means describing a measurement process in enough detail to be caught doing it.
This page is about publishing the experiment that produced nothing. Publishing figures only you have — the positive version of the same idea — is a different job, covered in how to create content AI engines actually cite.
What makes a negative result publishable?
A stated method, a stated sample, a named measurement surface, and stated limits. An absence with none of those is not a finding — it is a shrug with a headline, and readers can tell the difference immediately.
| Element | What it means | What happens without it |
|---|---|---|
| Method | Exactly what you deployed and where | "We tried it and nothing happened" is unverifiable |
| Sample and window | How many properties, over what dates | A one-week test reads as a law |
| Measurement surface | What would have shown an effect if one existed | The null cannot be distinguished from blindness |
| Limits | What the result does not prove | The page overclaims in the negative direction |
The fourth row is the one most often skipped, and it is the one that keeps a negative honest. Every operator-scale null contains some amount of missing instrumentation. If nothing in your stack reports the property you tested, a small effect could exist and be invisible to you, and the page has an obligation to say so rather than letting "we saw nothing" imply "there is nothing."
What do our own two negatives look like?
One is finished and one is still running, which is a useful contrast in itself.
| Experiment | Scope | Status | Result so far |
|---|---|---|---|
| Speakable schema | Every page of 3 production builds | Ongoing deployment, conclusion reached | No effect we can attribute to it [our data] |
| llms.txt watch | 3 production sites, 90-day window | In progress | 0 requests from documented AI engine crawlers so far [our data] |
| Ahrefs schema study (third party) | 1,885 pages | Published | No meaningful AI-citation lift from markup |
The finished one. We ship speakable markup on every page of three builds, pointing at the direct-answer and key-takeaways blocks, at a marginal cost of roughly zero because it is templated once. Nothing we track has moved in a way that points at the property — no citation observation, no referral pattern, no distinguishable crawl behavior [our data]. The full accounting, including the caveat that no tool in our stack reports the property at all, is in our speakable schema negative result.
The running one. We serve the llms.txt proposal (llmstxt.org) — not an adopted standard, and no major engine has committed to reading it — on three production sites and log every request to the path. The count so far is 0 requests from documented AI engine crawlers, on any of the three, while the same logs show those crawlers fetching HTML pages routinely [our data]. That window is still open and the page updates as it accrues; the running log is the llms.txt watch itself.
Publishing a count before the window closes is the harder version of this, and the better one. Declaring the sites, the window, and the counting method in advance means the result cannot be quietly reframed later depending on how it lands.
What does a null result actually prove?
That you did not observe an effect with the instruments you had, at the scale you ran. That is a real contribution and a narrow one, and conflating it with "the tactic does not work" is the mirror image of the overclaiming that makes this field untrustworthy in the first place.
Sample size decides how much weight a null carries. Three sites over weeks is an operator's honest maximum and a statistician's rounding error; Ahrefs' 1,885-page schema study is larger than any single agency's fleet and still observational. Our own reading of the markup evidence, ours and theirs together, is collected in what the evidence shows about schema and AI search.
Where a null is genuinely strong is against a specific promise. "Buy this and it will drive AI visibility" is refuted well enough by a fleet-wide deployment that produced no observable change, because the claim being sold was not subtle. A null is weak evidence about small effects and strong evidence against large advertised ones.
How do you write one without spin?
Lead with the result, not the setup. The first sentence states what you observed — "0 requests," "no effect we can attribute" — and the rest of the page earns it. Burying the null under three paragraphs of method is how a negative result gets quietly converted into a thought-leadership piece.
Four rules we hold ourselves to:
- Name the number, including when it is zero. "Minimal impact" is a hedge; "0 requests from documented AI engine crawlers so far" is a finding.
- Say what would have changed your mind. If a measurement surface exists that you did not have, name it.
- Keep the running experiment labeled as running. A window in progress is never reported as a concluded finding, however stable the number looks.
- Show where the public debate stands. On llms.txt, Google's John Mueller compared it to the keywords meta tag while Search Engine Land published a reasoned counterpoint. A zero is consistent with both positions, and saying so is more honest than recruiting the result to one side.
Google's people-first guidance frames the standard as demonstrated reliability and first-hand depth (Google, helpful content guidance). A documented failure is first-hand depth by definition — it cannot be written by anyone who did not run the thing.
When is publishing a negative the wrong move?
When you have no method to describe. An impression that something did not work, with no deployment record, no window, and no measurement surface, is not a negative result — it is an opinion wearing a lab coat, and publishing it adds another unverifiable claim to a field already drowning in them.
It is also wrong when the null is really a configuration error. Before publishing "X did nothing," confirm X was actually running: markup that failed validation, a file returning the wrong status, or a tag that never fired produces an absence that says nothing about the tactic.
Against our own interest, plainly: publishing negatives is not a growth tactic. It produces pages that end in "we could not attribute anything," which converts worse than confidence, costs the same to research as a success story, and occasionally argues a reader out of buying something we sell. The return is indirect and slow — it is that our positive claims become believable, because a library that publishes its nulls has demonstrated it is not selecting for the flattering ones. That trade is the whole editorial position behind this library, and it is a position, not a proven strategy.
Frequently asked questions
Should you publish experiments that didn't work?
Yes, if you can state the method, the sample, and the limits. A negative result is the one content type a competitor cannot fabricate cheaply, because faking a null requires describing a measurement process in enough detail to be caught.
What makes a negative result publishable?
Four things stated openly: what you did, on how many properties and for how long, what surface would have shown an effect, and what the result does not prove. Missing any of the four turns a finding into an unsupported absence.
Does a null result prove a tactic doesn't work?
No. It proves you did not observe an effect with the instruments you had, at the scale you ran. Part of every operator-scale null is missing instrumentation — if nothing in your stack reports the property, a small effect could exist unseen.
Is publishing failures good for business?
It is good for credibility and mixed for conversion. Pages that end in 'we could not attribute anything' convert worse than confident ones, and nobody should publish negatives expecting traffic. The return is that your positive claims become believable.
Can you publish a negative result before the experiment finishes?
Yes, if you label it as running. Our llms.txt watch publishes its count so far, names the window, and states the counting method before the result is complete — which is harder to walk back than a finding announced only once it looks good.
Sources
- Schema Markup and AI Citations: A 1,885-Page Study — Ahrefs
- The llms.txt proposal — llmstxt.org
- Google Says llms.txt Comparable To Keywords Meta Tag — Search Engine Journal
- No, llms.txt is not the new meta keywords — Search Engine Land
- Creating helpful, reliable, people-first content — Google