TL;DR
A randomized 2026 field experiment found 39.8% fewer outbound organic clicks when Google AI Overviews appeared and 34.5% more zero-click searches.
A separate three-experiment working paper found AI Overviews reduced source clicks, source-reading time and critical checking of product information.
Generative search engines also retrieve different source sets, so source selection is part of the answer, not just the prose.
A list of search results asks you to choose a source. AI search can give you a synthesized answer before you open one.
Saharsh Agarwal and Ananya Sen tested part of that change in a randomized field experiment using a Chrome extension. Some participants saw ordinary Google Search with AI Overviews. Others saw a version with the overview removed. Conditional on an overview appearing, outbound organic clicks fell 39.8% and zero-click searches increased 34.5%.
The paper is a working paper, not a peer-reviewed result. It also doesn’t show that people became less informed. It shows something narrower: when the answer appears before the source, people open fewer sources.
A second working paper posted September 3 gets closer to the mechanism. Matteo De Angelis, Marco Francesco Mazzù and Nicola Sabatini ran three online experiments with 584 participants using realistic clickable search environments. AI Overviews reduced source clicks and the time people spent reading sources during initial product search. A later experiment found lower epistemic vigilance, their term for critically checking information and its sources. That reduction indirectly increased purchase intention.
Again, product research isn’t every kind of search. But both studies found that putting a summary first reduced source clicks.
The Summary Comes First
Traditional search was never neutral. Ranking decided which links appeared first. Snippets framed what looked relevant. Advertising bought attention.
Search already offered answers without a click through featured snippets. Generative search can combine material from multiple sources before you open any of them.
Traditional search:
question -> ranked sources -> choose -> read -> synthesize -> judgment
Generative search:
question -> retrieve -> select -> synthesize -> answer -> optional sources -> judgmentThat convenience is real.
Before I see the answer, the engine may already have retrieved material, selected what to use and compressed disagreement into a few paragraphs.
The Consilience Project was writing about the larger sensemaking problem years before answer engines became normal. Its concern was that the information environment was growing more complex than our ability to make sense of it. A later Consilience essay argued that technologies aren’t neutral containers because the interface changes what receives attention and how people behave.
Generative search adds a specific mechanism to that older argument. The interface doesn’t only rank sources anymore. It can choose and combine them before you see them.
The Source Set Is Part of the Answer
A July 2026 paper in Findings of ACL compared Google organic search with five generative search systems from Google, OpenAI and Perplexity. The researchers found substantial differences in how the systems used internal versus external knowledge, which sources they retrieved and how stable those source sets were across executions. The systems could cover similar topics while pulling from different source sets and combining them differently. That means two fluent answers to the same question can be built from different evidence.
A September industry study from Prefer makes the variation easy to see, although I’d treat it as observational evidence rather than scientific proof. Prefer sent the same 80 questions to ChatGPT, Perplexity, Gemini and Claude three times each. Across 960 answers, 72.7% of the 1,329 cited sites appeared in only one engine. Only 2.2% appeared across all four. That’s a count of distinct sites, not citation volume. Among the 25 most-cited sites, 23 appeared in three or four engines.
The study used API access, so the results can differ from the consumer products. The question set was also about AI search, not the entire web.
Still, it illustrates the thing the ACL paper establishes more rigorously: a generative answer engine doesn’t simply summarize one fixed web. Which sources it retrieves helps determine the evidence behind the answer.
I wrote about a downstream version of this in Working Alongside: The Collaboration Asymmetry. That article was about AI shaping the frame of a decision after we ask for advice.
This starts earlier. Before the system frames what to do, it may already have shaped what evidence made it into view.
A Citation Can’t Show the Missing Sources
Citations help. I want more of them, not fewer. But a citation answers a limited question: where did this claim come from? It can’t tell you which credible sources the system retrieved and rejected, which ones it never retrieved or which disagreements disappeared during synthesis.
That distinction is easy to miss because citations make an answer look inspectable. Sometimes it is. But inspection starts with the selected source set.
In The Detector Is Not the Evidence, I wrote about the mistake of treating a confidence score as proof. This failure is different. Every visible citation can be accurate while important evidence is still missing.
A paper published September 25 in Ethics and Information Technology gives useful language for this. Qian Wu calls the set of sources allowed to count as grounds for a claim the evidence frame.
The paper is philosophical, not an experiment. I wouldn’t treat its framework as proof of user behavior. Its evidence frame concerns which sources are permitted for a task, not which ones an engine happens to retrieve. For AI search, my question is different: Which evidence did the system actually use?
Verification Has to Be an Action
A peer-reviewed study published in the Journal of Computer-Mediated Communication tested a conversational AI that delivered false health information to 477 participants. The researchers varied how conversational the assistant sounded and whether people had no verification option, saw a verification cue, were required to verify or could choose to verify.
Participants required to verify the claim rated the false information as less credible and trusted the assistant less than participants shown only the cue or no verification option. Seeing a disabled verification button didn’t have the same effect.
There’s an important limit here. The verification action sent participants to a page that debunked the false claim. This wasn’t a study of ordinary citation cards in AI search.
So I wouldn’t turn it into “citations don’t work.” My read is smaller: a route to the source and the act of checking the source are different things.
That fits the search studies. If the summary makes the click feel unnecessary, adding a source link can preserve the trail back to the source without getting anyone to follow it.
Compare the Evidence Before the Prose
The design question isn’t whether AI search should summarize information. That part is already useful. The question is how much of the evidence-selection process the interface leaves visible and easy to inspect.
For an answer that matters, I want four things:
Which source supports which claim?
Where do credible sources disagree?
Where is the evidence thin?
Can I get from the synthesis to the underlying material without fighting the interface?
There’s also a simple test anyone can run now. Ask the same consequential question in two different answer engines. Before comparing the prose, compare the sources. Open one source from each answer. Then look for one credible source neither answer cites.
If the cited source sets differ, the answers aren’t showing you the same evidence, even if they sound similar. That doesn’t make either answer wrong. It tells you something the fluent paragraph can hide: part of the answer was chosen before the writing began.
The source is becoming optional. The source selection isn’t.
Resources
Saharsh Agarwal and Ananya Sen, “The Impact of Google AI Overviews on Publisher Traffic and User Experience: Evidence from a Field Experiment”, working paper, posted April 3, 2026 and revised July 8, 2026.
Matteo De Angelis, Marco Francesco Mazzù and Nicola Sabatini, “The Summary Comes First: How AI Overviews Reshape Consumer Product Search”, working paper, posted September 3, 2026.
Elisabeth Kirsten, Jost Große Perdekamp, Qinyuan Wu, Mihir Upadhyay, Krishna P. Gummadi and Muhammad Bilal Zafar, “Characterizing Web Search in The Age of Generative AI”, Findings of ACL 2026, July 2026.
Mengqi Liao and S Shyam Sundar, “Chat but verify: Combating misinformation in conversational Generative AI with verification affordance”, Journal of Computer-Mediated Communication 31(3), published July 29, 2026.
Qian Wu, “From capability to assertability: epistemic regulation of LLM-mediated claim presentation”, Ethics and Information Technology 28, 46, published September 25, 2026.
The Consilience Project, “Challenges to Making Sense of the 21st Century”, March 30, 2021.
The Consilience Project, “Technology is Not Values Neutral: Ending the Reign of Nihilistic Design”, June 26, 2022.
Prefer, “How AI engines search and cite: 960 answers measured”, September 19, 2026. Industry observational study; used here as a current illustration rather than the article’s scientific backbone.
Evidence note: both Google AI Overviews behavioral studies are working papers rather than peer-reviewed findings. The 39.8% click effect is conditional on an AI Overview appearing. The De Angelis, Mazzù and Sabatini experiments concern product search. Prefer’s study is an industry observational dataset and used API responses that can differ from consumer applications.
Analytical note: the claim that answer engines construct an evidence environment before human judgment is my synthesis across the behavioral studies, retrieval research and Consilience’s sensemaking work. No single source measures that full concept directly.


