Research · Working Paper By Viveka Mohan Das 24 July 2026  ·  8 min read

I Couldn't Find the Research on the Perception Gap. So I Ran It.

Every AEO practitioner has watched it happen. A client's structural score climbs into the nineties in a matter of weeks. Then you ask ChatGPT or Gemini about the business, and the model draws a blank or invents something wrong. The client concludes the work failed. I've argued the opposite for a while, but a framework is only a story until someone measures it. When I went looking for that measurement, it didn't exist. So I built the study myself.

The framework itself is written up in plain English elsewhere: The Two-Clock Model explains the two clocks, the perception gap, and the three levers that close it. This piece is not that. This is the making-of: what it took to gather the evidence, why that rigour matters, how you can use the result with your own clients, and where the research goes next.

63.6pp
Mean structure-minus-perception gap at birth, across 48 entities. Zero counter-examples.
Dataset v1, frozen 8 Jul 2026
28 / 33
Entities where the citation ramp preceded AI perception onset. Median lead 83 days.
Two-sided sign test p = 6.6e-5
12 / 22
Audited entities clearing 70% citation precision. The rest collapse. Precision is bimodal.
566 hand-labelled articles
The birth gap — structure vs AI perception at entity launch (mean of 48 entities)
Structural presence AI perception
At birth, structure averages 64.2% of maximum while AI perception averages 0.5% — a 63.6-point gap with no counter-examples across all 48 entities measured both ways. By mid-2026, perception had overtaken structure for 32 of them. Data: dataset v1 (frozen 2026-07-08).

The hard part: you cannot observe the past of an AI model

The central problem is brutal. AI perception, what a model knew about an entity at a given moment, cannot be observed after the fact. Models are updated in place. Last year's version of ChatGPT is gone. There is no archive of what it used to believe.

The way around it is a natural experiment. Model vendors keep dated snapshots of older models, each with a stated training cutoff. If you ask a ladder of those snapshots the same question about the same entity, with browsing switched off, you are sampling what base models knew at a sequence of frozen moments in time. The curve of recognition against cutoff date is a reconstructed perception clock.

I ran 50 technology entities through five dated OpenAI snapshots, scored every answer 0 to 4 against ground truth, pulled archived homepages from the Wayback Machine at five points after each entity's birth, and built a weekly citation series from GDELT news mentions going back to each launch. That is the skeleton. The rigour is in the parts nobody sees.

The parts nobody sees

The concept, the analysis and the writing are my own, and I read and referenced every source by hand. An AI coding assistant helped write the data-collection scripts, and a language model acted as a scoring judge inside the experiment. Both are disclosed in full in the paper.

What the evidence actually shows

Three findings came out of the frozen dataset. I am stating them briefly, because the charts and the paper carry the detail.

1. The birth gap is universal

Across the 48 entities I could measure both ways, structure at birth averaged 64.2% of maximum while AI perception averaged 0.5%. That is a 63.6 percentage-point gap, with zero counter-examples. Every single entity showed structure ahead of perception at launch. By mid-2026, perception had overtaken structure for 32 of those 48. The gap is a phase, not a permanent state.

2. Mention-counting is bimodal, and that is a warning for the whole industry

Distinctive names measured above 70% precision. Common-word names collapsed, some to zero. If you are measuring a client's AI visibility by counting mentions, the entity's name alone decides whether your numbers mean anything. This is not a dial you tune with better queries. Either the name disambiguates or it does not, and you have to audit to know which.

Citation-query precision by entity — 566 hand-labelled articles, 22 entities
Pass — 70% or higher Fail — below 70%
Precision is not smoothly distributed — it is bimodal. Distinctive names disambiguate almost perfectly; common-word names collapse toward zero. Any dashboard that counts entity mentions inherits this problem. Data: dataset v1 (frozen 2026-07-08).

3. Citation tends to move first

Where I could observe both events, news-citation ramps preceded perception onset in 28 of 33 entities, a median of 83 days earlier. I report this as supporting evidence for the citation-transmission mechanism, not as proven cause. A common driver, genuine importance moving both coverage and model knowledge, would look the same. What the data rule out is the reverse: perception does not lead citation.

Days the citation ramp precedes AI perception onset — per entity, 33 entities
Citation leads (28) Perception leads (5)
Positive means citation moved first. The ramp precedes the onset in 28 of 33 entities (85%), median lead 83 days, two-sided exact sign test p = 6.6e-5. Data: dataset v1 (frozen 2026-07-08).

How other AEO agencies can use this

The perception gap is the hardest thing to explain to a client, because on paper the work looks done and the AI still can't name them. Until now, the only answer was "trust me, it's normal." That is a weak position to argue from, and clients can smell it. Now there is a citable, openly-licensed working paper behind the claim. You can use it three ways:

The paper is CC-BY-4.0. Cite it, quote it, hand it to a sceptical client. That is what it is for.

This is a working paper. Here is the next harvest.

I am calling this v1 for a reason. The natural experiment that made it possible is perishable: the oldest snapshots that anchor the early end of the perception curve are scheduled to retire through late 2026. Once they are gone, that measurement cannot be rebuilt. So the next harvest is time-boxed, and it is soon.

WindowMilestoneWhat changes
Aug 2026 v2 design freeze Lock a repeat-run protocol (3 passes per cell) so perception scores carry variance and inter-rater reliability, not a single reading.
Sep–Oct 2026 v2 harvest (before snapshots retire) Re-run every perception cell 3 times; add a second and third vendor's dated snapshots while they still exist; re-query the 10 failed-precision entities with tightened disambiguation.
Nov 2026 v2 freeze + analysis Report variance bands on the perception curve, inter-rater reliability, and a refreshed precedence test on the audited-only subset.
2027 v3 scope Extend the panel beyond AI-sector entities and English-only coverage; run a formal Granger causality test once perception resolves at finer than four dates.

Where this is going, and why I'm inviting others in

The timetable above is the near term. The longer arc is more ambitious, and I will be honest about why it needs more hands than mine. Three questions decide whether the two-clock model becomes a durable piece of AEO knowledge or stays a promising one-off:

Here is the honest part. I am running this as one person, around a full-time role and a full client load. That is the real constraint on how fast v2 and v3 arrive. It is also exactly why I froze the data, published the code, and licensed the whole thing openly: so the work does not live or die by my calendar.

So this is an open invitation. If you are a researcher, an AEO or GEO practitioner sitting on an entity panel, someone with cross-vendor snapshot access, a statistician who works in causal inference, or simply someone who wants to blind-relabel a precision subset and try to break my numbers, I would genuinely like to hear from you. Co-authorship is on the table for substantive contributions. The dataset is public, the limitations are documented, and the next harvest is scoped and waiting.

Read it, check it, replicate it

The full working paper and the complete dataset, code and scripts are published openly on Zenodo under a CC-BY-4.0 licence:

The dataset includes the entity roster with sourced birth dates, all perception and structure scores, the 6,570-row frozen citation series with per-entity precision labels, and the Python collection and analysis scripts. Anyone can download it, regenerate the figures, and try to break the result. That is the point of publishing it. For the framework in plain English, start with The Two-Clock Model, and see it running live on the Client Zero Visibility Dashboard.

Want to collaborate — or see where your business sits?

If you'd like to contribute to v2, or you want an AI Visibility Audit that measures your structural presence and your current standing in AI answers, get in touch.

Get in Touch    Check Your AEO Score Free

Working paper, not peer-reviewed. Competing interest: the author is the founder of AISearch Global, which sells answer engine optimisation and generative engine optimisation services. The two-clock model is used in that practice, and that conflict is disclosed throughout the paper. Author ORCID: 0009-0007-9715-0951.