Where AI answers get their sources: what 680 million citations show
A synthesis of the largest 2025–2026 citation datasets — Profound, Goodie, Surfer, Semrush, Peec, Ahrefs — on which domains ChatGPT, Perplexity, and Google AI Overviews actually cite, how little the platforms overlap, and what a single site can do about it.
The three big answer engines cite from three different webs. Wikipedia accounts for 47.9% of ChatGPT’s top-10 source share, Reddit for 46.7% of Perplexity’s, and Google AI Overviews cite at least one top-10 organic result in 93.67% of responses. Only about 11% of domains are cited by both ChatGPT and Perplexity, and only 12% of URLs cited by the AI tools overlap with Google’s top ten. Those figures come from 5W’s June 2026 “State of AI Citations” report, which synthesized nine public datasets covering more than 680 million tracked citations.
This article lays out the numbers and what they mean for a site that is not Wikipedia or Reddit.
What data is this built on?
5W combined citation datasets published between August 2024 and April 2026:
| Source | Scope |
|---|---|
| Profound | 680 million citations across ChatGPT, Google AI Overviews, and Perplexity, August 2024–June 2025 |
| Goodie | 5.7 million citations (Feb–Jun 2025), expanded to 58.6 million (Oct 2025–Mar 2026) |
| Surfer AI Tracker | 46 million citations across 36 million AI Overviews, March–August 2025 |
| Semrush | 230,000 prompts across ChatGPT, Google AI Mode, and Perplexity over thirteen weeks |
| Peec AI | 30 million citations, March 2026 snapshot |
| Ahrefs | 15,000-query comparison of AI citations against organic rankings |
| BrightEdge and WebFX | Healthcare vertical, 130,000+ queries |
The datasets use different prompt sets and time windows, so the synthesized shares are estimates, not a single measurement. The direction of every finding below is consistent across the underlying sources.
Which domains dominate?
ChatGPT leans on reference sites. Wikipedia takes 47.9% of ChatGPT’s top-10 source share and 7.8% of all ChatGPT citations across the full dataset. ChatGPT also mentions brands about 3.2 times more often than it links to them, so a brand can be present in answers without a single citation showing up in a tracker.
Perplexity leans on community and review sites. Reddit takes 46.7% of Perplexity’s top-10 share, and the rest of its top tier is G2, Gartner, NerdWallet, PCMag, TripAdvisor, and Yelp. Perplexity is the engine most driven by third-party opinion.
Google cites Google, and then its own top ten. Roughly 43% of AI Overview citations link to Google-owned properties. Of the rest, 93.67% of responses include at least one top-10 organic result, yet only 4.5% of Overview URLs directly match a page-one organic URL. Google cites the ranking domain’s deeper page, not the ranking page.
How stable are these shares?
Not very. The report documents a September 2025 swing in which Reddit’s share of ChatGPT citations fell from roughly 60% to 10% and Wikipedia’s presence dropped from about 55% of responses to under 20%, in the space of weeks, with no announcement. Any strategy built on one month’s leaderboard is fragile; any strategy built on being fetchable and extractable is not.
Does ranking on Google predict AI citations?
Only on Google. Across ChatGPT and Perplexity, 88% of cited URLs do not rank on Google’s page one. The strongest single predictor the report identifies is not rank at all but brand search volume, with a 0.334 correlation to citation likelihood. Engines cite what people already search for by name.
How much traffic is left?
Zero-click searches rose from 56% of queries in 2024 to 69% by May 2025 in the report’s data, and referral traffic to news publishers fell from 2.3 billion monthly visits in mid-2024 to under 1.7 billion by May 2025. Citation is increasingly the whole prize; the click that used to follow it is optional.
What does a single site do with this?
The report is about the web as a whole. Translated to one domain, five things follow.
1. Fetchability and extractability are the on-page ceiling. Every dataset counts citations to pages an engine could retrieve and parse. Nothing else in this article matters for a page that returns a bot challenge or renders client-side. That is why the audit puts 43 of its raw weight on fetchability.
2. Be citable at the paragraph level, not the page level. The 4.5% URL-match figure means engines pick the tightest answer on a domain, not its most authoritative page. Each page should carry one clear question and a direct answer in its first 200 words; the front-loaded answer check exists for this.
3. Your off-site presence is half the game, and no page audit can see it. Reddit threads, Wikipedia references, G2 and Yelp listings, and news mentions are where two of the three engines get most of their sources. For a local business or a niche product, an accurate presence on the review site an engine already trusts may earn more citations than any change to your own site.
4. Brand demand is the strongest predictor, so measure mentions, not just links. ChatGPT’s 3.2:1 mention-to-link ratio means link-based tracking undercounts you by a wide margin.
5. Plan per engine. With 11% overlap between ChatGPT and Perplexity, a page that wins on one has no claim on the other. Reference-quality explainers with clean structure travel best on ChatGPT and AI Mode; community and review presence travels best on Perplexity.
What the audit measures, and what it cannot
The score covers the on-page half: whether an engine can fetch the page, whether the article body is extractable, whether the entity is identified in JSON-LD, whether the answer is front-loaded. It does not and cannot measure Reddit share of voice, brand search volume, or whether any engine has cited you. Treat a high score as a necessary condition, and the numbers above as the reason it is not a sufficient one.
Further reading
- The Princeton GEO paper, summarized
- Google AI Overviews and AI Mode: what actually decides whether you’re in
- AEO vs SEO: what’s actually different
Check the on-page half in seconds: Run an audit →