I Couldn't Find the Research on the Perception Gap. So I Ran It.
Every AEO practitioner has watched it happen. A client's structural score climbs into the nineties in a matter of weeks. Then you ask ChatGPT or Gemini about the business, and the model draws a blank or invents something wrong. The client concludes the work failed. I've argued the opposite for a while, but a framework is only a story until someone measures it. When I went looking for that measurement, it didn't exist. So I built the study myself.
The framework itself is written up in plain English elsewhere: The Two-Clock Model explains the two clocks, the perception gap, and the three levers that close it. This piece is not that. This is the making-of: what it took to gather the evidence, why that rigour matters, how you can use the result with your own clients, and where the research goes next.
The hard part: you cannot observe the past of an AI model
The central problem is brutal. AI perception, what a model knew about an entity at a given moment, cannot be observed after the fact. Models are updated in place. Last year's version of ChatGPT is gone. There is no archive of what it used to believe.
The way around it is a natural experiment. Model vendors keep dated snapshots of older models, each with a stated training cutoff. If you ask a ladder of those snapshots the same question about the same entity, with browsing switched off, you are sampling what base models knew at a sequence of frozen moments in time. The curve of recognition against cutoff date is a reconstructed perception clock.
I ran 50 technology entities through five dated OpenAI snapshots, scored every answer 0 to 4 against ground truth, pulled archived homepages from the Wayback Machine at five points after each entity's birth, and built a weekly citation series from GDELT news mentions going back to each launch. That is the skeleton. The rigour is in the parts nobody sees.
The parts nobody sees
- 566 articles, labelled by hand. Entity names are search queries, and search queries lie. "Mamba" returns Skoda cars. "Operator" returns tour operators. "Dream Machine" returned an obituary cycle. I hand-labelled 566 sampled articles across 22 entities as relevant or not, because a citation count you have not audited is a number you cannot trust.
- The dataset was frozen before analysis. On 8 July 2026 I locked the data — 6,570 entity-weeks of citations plus every perception and structure score — and gave it a version number. Nothing changed after. That is the difference between finding a pattern and fishing for one.
- I documented what I did not do. The perception scores are a single-run pilot, so I do not yet have run-to-run variance. Citation precision is bimodal, so I treat only clean-name entities as high-confidence. I sell this service, so that conflict is disclosed on every version. A working paper that hides its limitations is marketing; one that states them plainly is research.
The concept, the analysis and the writing are my own, and I read and referenced every source by hand. An AI coding assistant helped write the data-collection scripts, and a language model acted as a scoring judge inside the experiment. Both are disclosed in full in the paper.
What the evidence actually shows
Three findings came out of the frozen dataset. I am stating them briefly, because the charts and the paper carry the detail.
1. The birth gap is universal
Across the 48 entities I could measure both ways, structure at birth averaged 64.2% of maximum while AI perception averaged 0.5%. That is a 63.6 percentage-point gap, with zero counter-examples. Every single entity showed structure ahead of perception at launch. By mid-2026, perception had overtaken structure for 32 of those 48. The gap is a phase, not a permanent state.
2. Mention-counting is bimodal, and that is a warning for the whole industry
Distinctive names measured above 70% precision. Common-word names collapsed, some to zero. If you are measuring a client's AI visibility by counting mentions, the entity's name alone decides whether your numbers mean anything. This is not a dial you tune with better queries. Either the name disambiguates or it does not, and you have to audit to know which.
3. Citation tends to move first
Where I could observe both events, news-citation ramps preceded perception onset in 28 of 33 entities, a median of 83 days earlier. I report this as supporting evidence for the citation-transmission mechanism, not as proven cause. A common driver, genuine importance moving both coverage and model knowledge, would look the same. What the data rule out is the reverse: perception does not lead citation.
How other AEO agencies can use this
The perception gap is the hardest thing to explain to a client, because on paper the work looks done and the AI still can't name them. Until now, the only answer was "trust me, it's normal." That is a weak position to argue from, and clients can smell it. Now there is a citable, openly-licensed working paper behind the claim. You can use it three ways:
- To reset expectations at the start. Show a client the birth-gap number. A brand-new or newly-optimised entity sits, on average, 63 points ahead on structure and near zero on perception. That is not your failure. That is the starting line for everyone.
- To justify the citation-building phase. The finding that citation ramps precede perception gives you the argument for why the budget shifts from schema to external mentions once the structural work is done. It is the lever that moves the slow clock.
- To defend your measurement. If you report AI-visibility numbers, the bimodal-precision result tells you which of a client's entity names you can actually trust a mention count for, and which need a disambiguation audit first.
The paper is CC-BY-4.0. Cite it, quote it, hand it to a sceptical client. That is what it is for.
This is a working paper. Here is the next harvest.
I am calling this v1 for a reason. The natural experiment that made it possible is perishable: the oldest snapshots that anchor the early end of the perception curve are scheduled to retire through late 2026. Once they are gone, that measurement cannot be rebuilt. So the next harvest is time-boxed, and it is soon.
| Window | Milestone | What changes |
|---|---|---|
| Aug 2026 | v2 design freeze | Lock a repeat-run protocol (3 passes per cell) so perception scores carry variance and inter-rater reliability, not a single reading. |
| Sep–Oct 2026 | v2 harvest (before snapshots retire) | Re-run every perception cell 3 times; add a second and third vendor's dated snapshots while they still exist; re-query the 10 failed-precision entities with tightened disambiguation. |
| Nov 2026 | v2 freeze + analysis | Report variance bands on the perception curve, inter-rater reliability, and a refreshed precedence test on the audited-only subset. |
| 2027 | v3 scope | Extend the panel beyond AI-sector entities and English-only coverage; run a formal Granger causality test once perception resolves at finer than four dates. |
Where this is going, and why I'm inviting others in
The timetable above is the near term. The longer arc is more ambitious, and I will be honest about why it needs more hands than mine. Three questions decide whether the two-clock model becomes a durable piece of AEO knowledge or stays a promising one-off:
- Is it real beyond one vendor? v1 reconstructs perception from OpenAI snapshots alone. If the same birth gap and the same citation-first ordering show up in Anthropic and Google snapshot ladders, the finding hardens from an OpenAI artefact into a property of how language models learn entities. This is the single most valuable replication, and the window to run it is closing.
- Is it cause, or coincidence? Right now I can show citation moves before perception. I cannot yet show citation moves perception. Closing that gap needs finer-grained perception measurement and a formal causality test, a heavier statistical lift than a solo pilot can carry well.
- Does it hold outside tech? The panel is AI-sector entities with dense English coverage, the easy case. The businesses this framework is meant to help are small and mid-market firms in ordinary sectors, where citation is sparser and slower. Whether the gap dynamics generalise there is the question that matters most for practitioners, and it is completely untested.
Here is the honest part. I am running this as one person, around a full-time role and a full client load. That is the real constraint on how fast v2 and v3 arrive. It is also exactly why I froze the data, published the code, and licensed the whole thing openly: so the work does not live or die by my calendar.
So this is an open invitation. If you are a researcher, an AEO or GEO practitioner sitting on an entity panel, someone with cross-vendor snapshot access, a statistician who works in causal inference, or simply someone who wants to blind-relabel a precision subset and try to break my numbers, I would genuinely like to hear from you. Co-authorship is on the table for substantive contributions. The dataset is public, the limitations are documented, and the next harvest is scoped and waiting.
Read it, check it, replicate it
The full working paper and the complete dataset, code and scripts are published openly on Zenodo under a CC-BY-4.0 licence:
- Whitepaper (Zenodo)https://doi.org/10.5281/zenodo.21533068
- Dataset and code (Zenodo)https://doi.org/10.5281/zenodo.21532575
The dataset includes the entity roster with sourced birth dates, all perception and structure scores, the 6,570-row frozen citation series with per-entity precision labels, and the Python collection and analysis scripts. Anyone can download it, regenerate the figures, and try to break the result. That is the point of publishing it. For the framework in plain English, start with The Two-Clock Model, and see it running live on the Client Zero Visibility Dashboard.
Want to collaborate — or see where your business sits?
If you'd like to contribute to v2, or you want an AI Visibility Audit that measures your structural presence and your current standing in AI answers, get in touch.
Get in Touch Check Your AEO Score FreeWorking paper, not peer-reviewed. Competing interest: the author is the founder of AISearch Global, which sells answer engine optimisation and generative engine optimisation services. The two-clock model is used in that practice, and that conflict is disclosed throughout the paper. Author ORCID: 0009-0007-9715-0951.