SEO

Join 500+ brands growing with Passionfruit!
A marketing team ships a set of structural changes, waits sixty days, and finds ChatGPT referrals up 5.7x. The obvious conclusion is that the work succeeded.
That exact result appears in a field experiment published as an arXiv preprint, "Disentangling AEO from Platform Growth." The researchers ran the intervention on a live site and tracked the intervened pages over several months. Referrals grew 5.7x.
They also tracked a control group: pages on the same site that received no changes at all. Those grew 3.5x.
Once platform growth was isolated, the net effect of the intervention was roughly 1.82x, with a 95% confidence interval of 1.31 to 2.54. Real, meaningful, and about a third of the headline number.
Almost every published before-and-after AEO result omits that second measurement. Which means almost every published before-and-after AEO result is reporting platform growth plus intervention effect, and calling the total intervention effect.
This piece covers the study design that separates the two, why the existing published research cannot answer the question most teams are actually asking, and how to run the measurement on your own domain.
Why the Existing Research Cannot Answer Your Question
The published work on AI citations in 2026 is substantial and mostly well-executed. It is also almost entirely one study type.
What a census measures
A census samples many domains at one point in time and reports which content gets cited most. The major ones:
Wix Studio's AI Search Lab and HubSpot's State of AEO, reported via Search Engine Land, analysed over one million citations between them. Listicles accounted for 21.9% of citations, articles 16.7%, and product pages 13.7%, together more than half of all citations. Comparison content held the highest single-citation rate on ChatGPT of any format.
Vemetric tracked 768,000 citations and found product-related content accounting for 46% to 70% of cited sources, with news and research articles at 5% to 16%.
BrightEdge, via Averi's benchmarks, analysed more than 50,000 AI-generated responses and found original research cited at roughly 20 times the rate of thin content and 4 to 10 times the rate of a standard blog post.
MaxAEO logged 61,400 citations from 9,200 B2B and SaaS-intent prompts over eight weeks, finding original-data studies at a 71% citation rate, comparison pages at 64%, and ranked listicles at 61%.
TripleDart found a 30x to 50x citation gap between the highest and lowest performing page types on identical domains.
This is good data. It tells you what correlates with citation across the open web, and the consistency across independent samples makes the direction trustworthy.
What a census cannot tell you?
It cannot tell you what will happen if you change something.
A census is observational and cross-sectional. It reports that pages with certain properties get cited more often, measured across thousands of domains you do not control, at a moment that has already passed. It does not isolate cause; it does not account for the confound that better-resourced domains produce both better-structured pages and more citations for unrelated reasons, and it cannot tell you the size of the effect on a domain with your starting authority.
That last point is not a technicality. SE Ranking's analysis of 2.3 million pages found domain traffic to be the single strongest predictor of AI citation, with high-traffic sites earning approximately three times more citations than low-traffic ones. If authority dominates, then a format finding drawn mostly from high-authority domains may not transfer to yours at anything like the same magnitude.
What an intervention study measures instead
An intervention study inverts the design. Instead of many domains once, it measures one domain repeatedly: log every change shipped, with dates and URLs, track citation rate across runs, and hold back a matched set of pages that receive nothing.
It trades external validity for internal validity. The findings apply to one domain, but within that domain you can actually attribute movement. For a team deciding where to spend next quarter's editorial budget, that is the more useful trade.
The Measurement Problems That Make This Harder Than It Looks
Four things break naive before-and-after measurement in this category. All four are documented.
Citation is non-deterministic
The same prompt, run twice in the same hour, returns different sources. SE Ranking's 10,000-keyword study found Google's AI Mode returning roughly 91% different URLs across three same-day repeat searches of the same query, with only 9.2% URL overlap. AirOps found approximately 30% of brands remaining visible across consecutive AI answers for the same query.
A single check is one draw from a probability distribution. If you ran a prompt before your changes and again after, and your page appeared the second time, you have learned almost nothing.
Implication: run every prompt at least three times per measurement cycle and report the average. More runs if the budget allows.
Platforms are not interchangeable.
Profound reported ChatGPT and Claude sharing only 8% domain overlap, and separately that 79.2% of Claude citations come directly from Brave's top 10 search results. SE Ranking found AI Overviews and AI Mode sharing just 10.7% of cited URLs and 16% of domains, despite both being Google products.
These are different retrieval systems producing similar-looking output. A blended "AI visibility score" averaged across them can rise while your visibility on the platform your buyers actually use is falling.
Implication: report citation rate per platform, never averaged. A composite number is a reporting convenience, not a measurement.
Ranking and citation have decoupled
Ahrefs' study of 15,000 long-tail queries found only 12% overlap between AI assistant citations and Google's top 10. Roughly 80% of URLs cited across ChatGPT, Perplexity, Copilot, and AI Mode do not rank in Google's top 100 for the original query.
Implication: rank is not a proxy for citation and cannot stand in for it in a measurement design. You have to measure citation directly.
The platforms are growing underneath your data
This is the one that invalidates most published results. AI referral traffic and AI citation volume are both rising fast on their own. Any page, changed or unchanged, is likely to gain citations across a sixty- or ninety-day window simply because more queries are being run and more citations are being issued.
The arXiv field experiment quantified it: 3.5x growth on pages that received no changes whatsoever.
Implication: a control set is not optional rigour. Without one, you cannot distinguish your work from the tide.
The Study Design That Actually Isolates Effect
Six steps. None require a tool you do not already have, though the repetition is tedious by hand.
1. Define the prompt set and freeze it
Pick 30 to 50 buyer-intent prompts that reflect how your buyers actually phrase problems, not your keyword list. Freeze the set for the duration of the study. Adding prompts mid-study makes the periods incomparable.
2. Establish a baseline with repetition
Run every prompt at least three times per cycle, across every platform you care about, before shipping any changes. Record which of your URLs were cited, on which platform, in response to which prompt. Average across runs. This is your baseline, and you cannot reconstruct it later.
3. Split intervened and control sets
Match pages on type and starting position, then assign roughly half to receive changes and half to receive nothing for the full window. The control set must be genuinely untouched, including content refreshes and internal linking, or it stops functioning as a control.
The temptation to skip this is strong, because the control pages are pages you are choosing not to improve. That cost is the price of knowing whether the improvements work.
4. Log every change with dates and URLs
Schema added, entity and author markup, headings restructured, content refreshed, internal links added, rendering fixes. Each one dated, each one tied to specific URLs. Without this log, you can measure that something moved but not what moved it.
5. Re-measure on the same cadence
Same prompts, same run count, same platforms, same intervals. Consistency of method matters more than frequency.
6. Report the net, and report the nulls
State the intervened movement, the control movement, and the difference. If the control rose by most of what the intervened set rose, say so. And report the changes that produced nothing alongside the ones that worked, with the same prominence.
Null results are the most useful output of this design and the part the published literature rarely includes. A reader learns more from "we added schema to forty pages and citation rate did not move" than from another confirmation that original research gets cited.
What This Design Still Cannot Prove
Being honest about the limits is part of why the design is worth using.
One domain is one domain. A control set removes the platform-growth confound. It does not make your results generalisable. Your starting authority, category, and existing citation base all shape the magnitude, and SE Ranking's 2.3 million page analysis suggests authority dominates.
Model updates remain a confound. Retrieval behaviour changes underneath you during any measurement window, and a control set only partly absorbs that. If a platform changes how it selects sources mid-study, both sets move for reasons unrelated to your work.
Owned content has a ceiling. Omniscient Digital's analysis of 23,387 citations using Peec AI found that on branded queries, earned media accounts for 48% of citations, commercial brand content for 30%, and owned brand content for just 23%. On-site structural work improves the share you control, which is roughly a quarter of the citations that mention you. The rest lives on review sites, forums, and press, and no amount of schema markup reaches it.
Sample sizes in this design are small. You are working with tens of prompts, not a million citations. The design buys attribution, not statistical power. Treat the output as a directional read on your own domain rather than a finding about the category.
What to Do With This on Monday
Check crawler access before anything else. No major AI crawler executes JavaScript. GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended read server-rendered HTML only. Confirm your content is in the server response, and check robots.txt for each crawler individually, since permitting one does not affect the others. A rendering defect makes every other optimisation irrelevant.
Build the baseline this week. Thirty prompts, three runs each, recorded per platform. It is a few hours of work, and it is the only thing you cannot backfill.
Pick your control set before you are tempted not to. Match on page type and starting position, write the list down, and leave those pages alone.
Stop reporting a blended AI visibility number. Split it by platform. If the split has never been shown to your leadership, showing it is usually the most informative single slide in the deck.
Ask any vendor or agency showing you a before-and-after what their control was. If there was not one, the number includes platform growth, and you have no way to know how much.
If you want the per-platform view without building the tracking by hand, Passionfruit Labs runs a fixed prompt set and reports citation share by platform. Our AI visibility tracking page covers how the measurement works, and page-level analytics covers the per-URL view this design depends on.
How Passionfruit Approaches AI Citation Work
Passionfruit is a managed SEO and AEO partner for DTC and B2B SaaS brands. We run tracking, analysis, and execution as one motion rather than handing over a dashboard, which is what makes this study design usable in practice: the team logging the citation data is the team shipping the changes, so the intervention log and the measurement come from the same place.
We have a commercial interest in this argument. We sell the work this design measures, and a methodology that makes AEO effects measurable is a methodology that helps us prove our own value. We are publishing the design, including the parts that shrink headline numbers, so readers can weigh that interest with full information. Our Search Console data-integrity research carried the same disclosure.
We are running this design on our own domain now. Fixed prompt set, per-platform reporting, matched control group, full intervention log. We will publish the results, including the control movement and the null results, whatever they show. Our research hub carries our other first-party analysis in the meantime, including our study of Search Console impression inflation across 17 properties, our examination of why AI citations may not be the visibility metric worth tracking, and our scoring of 350 top-ranking pages for AI content signals.
To discuss running this on your own domain, explore our generative engine optimization service or talk to an expert.
Frequently Asked Questions
What is a control set in AI citation measurement?
A group of pages on your own site, matched to your intervened pages on type and starting position, that receive no changes during the measurement window. Their citation movement shows you how much your intervened pages would have moved anyway. The arXiv field experiment "Disentangling AEO from Platform Growth" found control pages growing 3.5x while intervened pages grew 5.7x, leaving a net intervention effect of roughly 1.82x.
Why can't I just compare my citation rate before and after?
Because AI platforms are growing fast enough that untouched pages gain citations on their own. A before-and-after without a control reports platform growth plus intervention effect as a single number, and you have no way to know the split. In the one published experiment that measured both, platform growth accounted for the majority of the raw increase.
How many times should I run a prompt before treating the result as data?
At least three per measurement cycle. Google's AI Mode returns roughly 91% different URLs across three same-day repeat searches of the same query, and only about 30% of brands stay visible across consecutive AI answers. One check samples a probability distribution once and tells you very little.
Should I report one AI visibility score or split it by platform?
Split it. ChatGPT and Claude share only 8% domain overlap in Profound's data, and Google's own AI Overviews and AI Mode share just 10.7% of cited URLs. A blended score can rise while visibility falls on the platform that matters to your buyers.
Do AI crawlers read JavaScript-rendered content?
No. GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, and Google-Extended read server-rendered HTML only. Client-side content is absent from the candidate set regardless of how the page ranks in Google. Check each crawler separately in robots.txt, since blocking or permitting one does not affect the others.
Can on-site changes alone move a brand's AI citation rate?
Partly, with a ceiling. Omniscient Digital's analysis of 23,387 citations found earned media at 48% of citations on branded queries against 23% for owned brand content. Structural work improves the share you control. The rest depends on presence in sources you do not own.





