Proof
What didn't work
Interventions we've run that produced no measurable change, or made things worse. Dated, with the data. We publish these on the same cadence as the results that worked, because an agency that only shows you its wins is showing you a marketing asset rather than a track record.
Last updated 31 July 2026
Why this page exists
Ask any AEO agency to show you an intervention that failed. Almost none can, and the reason isn't that they don't have failures, it's that publishing them is commercially uncomfortable.
But consider what a failure log actually proves. It proves the measurement is real, because you can't detect a null result without a method sensitive enough to distinguish one. It proves the reporting is honest, because nobody fabricates a disappointment. And it proves the practitioner understands their own field, because in a domain this young a competent operator should expect roughly half their hypotheses to fail.
This page is also the single hardest thing for a competitor to copy. They'd have to have been measuring properly for long enough to have real nulls, and then be willing to publish them.
How we classify a null result
No change detected, the post-intervention confidence interval overlaps the pre-intervention interval. We cannot distinguish the result from noise. This is not the same as "it did nothing"; it means our measurement, at the depth we ran it, could not detect an effect.
Decline, the interval moved down and does not overlap. Something got worse, though not necessarily because of us.
Confounded, a platform change, a competitor action or an external event occurred in the measurement window and we cannot attribute the movement either way. We publish these rather than quietly discarding them, because discarding inconvenient windows is how honest reporting turns into selective reporting.
Entry template
Every figure below is EXAMPLE. This block shows format only.
[Short description of the intervention]
Client: [name or anonymised, sector] · Run: [date] · Measured over: [n] cycles
What we did and why we expected it to work:
The hypothesis, stated as it was before the result was known. Be specific about what you predicted, retrospectively vague predictions are unfalsifiable.
What happened:
| Before | After | Verdict | |
|---|---|---|---|
| ChatGPT SOV | EXAMPLE 6.1% ±1.2 | EXAMPLE 6.4% ±1.2 | Intervals overlap, no change detected |
What we think happened:
Honest analysis. "We don't know" is an acceptable answer and appears here more often than anywhere else on the site.
What we changed as a result:
Whether it's off the deliverables list, being retested differently, or retained on other grounds.
Published nulls
Empty until we have real ones. This page is deliberately not seeded with plausible-sounding examples, a fabricated failure is a worse offence than a fabricated success.
First entries expected [DATE], following the completion of the first full measurement cycles.
In the meantime, here's what we already decline to sell on the basis of other people's published evidence rather than our own:
llms.txt. Cyrus Shepard's 2026 synthesis scored 23 citation factors across 54 studies, patents and experiments. llms.txt came last at 2.0 out of 10, with no credible evidence of measurable effect in any source. Google has said it won't use it. We haven't tested it ourselves because the existing evidence base is sufficient and running the test would mean charging a client for it.
Content chunking for hypothetical retrieval models. No published evidence, and the implementation details it depends on are undocumented and change without notice.
Bought brand mentions. We won't test this one on ethical grounds rather than evidential ones.