Signal Brief #003
The Citation Oligarchy
When an AI answers a question, it does not survey the web. It quotes a handful of the same domains over and over. The citation layer of the internet is collapsing onto a few hubs, and the concentration is more extreme than search rankings ever were.
The last brief made a narrow point: a top Google ranking no longer predicts whether an AI will cite you. That raises an obvious follow-up. If not the ranked pages, then what does the machine quote?
The answer is unsettling in a different way. It quotes the same few websites, almost regardless of the question.
The shape of the citation layer
When researchers look at where AI answers actually pull their sources from, the same names surface at the top of every list. Peec AI analyzed roughly 30 million cited sources across the five major engines (ChatGPT, Google AI Mode, Gemini, Perplexity, and AI Overviews). The same names sit at the top: Reddit first, then YouTube, LinkedIn, Wikipedia, and Forbes.
Most of these are not subject-matter authorities. They are hubs: high-volume, densely linked, endlessly crawled platforms that sit near the center of the web’s link graph. The machine is not finding the best source on your question. It is reaching for the nearest large, reachable, well-connected node and quoting that.
By one synthesis of six citation studies (roughly 680 million citations in total), the top 15 domains capture about 68% of all AI citation share – a concentration the authors call more extreme than Google PageRank ever produced. Treat the exact figure as a vendor estimate; the pattern it names is corroborated everywhere you look.
More concentrated than search ever was
Search has always had popular domains, but a results page still offered ten blue links and a long tail beneath them. The citation layer does not. An answer quotes two or three sources, so the act of answering is itself a concentration mechanism: every reply is a vote, and the votes pile onto the same handful of hubs.
That is the structural difference worth sitting with. A search engine distributes attention across a ranked list a person can scroll. An answer engine collapses attention to a point. When the unit of output shrinks from a page of links to a single sourced paragraph, the distribution of who gets cited stops looking like a ranking and starts looking like a power law: a few enormous winners, and a very long tail of sources that are quoted essentially never.
And your own site lives in that tail. In one category measurement of answer-engine citations, owned media (a brand citing its own domain) accounted for under 1.5% of citations. The pages you control are not where the machine looks. It looks at the hubs that talk about you.
Each engine has its own oligarchy
The concentration is real, but it is not one shared list. Each engine funnels toward a different set of hubs, shaped by how it was built and what it was trained or paid to trust.
ChatGPT leans on Wikipedia, Reddit, and editorial sites like Forbes. Perplexity still concentrates on Reddit, YouTube, LinkedIn, Wikipedia, and B2B review sites like G2. Google’s AI Mode tilts toward its own orbit, surfacing Facebook and Yelp. Gemini sits closer to Reddit, YouTube, and Wikipedia. A separate synthesis that includes Claude finds a literary tilt: the New York Times, the Atlantic, the New Yorker, and the Economist. Same web, different small worlds.
For anyone trying to “show up in AI,” this is the part that breaks the SEO instinct. There is no single list to climb. Being the kind of source Perplexity trusts looks nothing like being the kind of source ChatGPT reaches for.
The hubs are not even stable
The last reason chasing a specific domain is a trap: the hubs move. The same synthesis that measured the concentration also clocked ChatGPT’s Reddit citation share falling from roughly 60% to 10% in about two weeks in late 2025. A model update, a licensing deal, a safety tweak, and the source mix lurches. The concentration is durable; the identity of the winners is not.
So a strategy of “get cited on Reddit” or “win Wikipedia” is building on ground that shifts in weeks. You can chase the current hub and find it demoted by the next model release.
What this does not say
Stated plainly, so the framing does not outrun the data.
The headline 68% is a vendor synthesis, not a single controlled study, and citation shares are measured differently across the six it folds together. The per-engine lists come from snapshots that themselves drift. And concentration is not the whole story: long-tail and primary sources do get cited, especially by retrieval-grounded engines on specific technical questions. The claim that survives every caveat is the structural one. The citation layer is far more concentrated than the search results page it is replacing, the winners are a small set of hubs, and those hubs are both engine-specific and unstable.
Why this is a structure problem, not a PR one
Here is the through-line from the rest of this series, and the part we measure for a living.
If the machine mostly quotes hubs, the question is not “how do we become a hub” (you almost certainly will not) but “what does a hub find when it looks at us, and what can a retrieval pass reach on our own site when it does come?” Both are structural. A Reddit thread or a Wikipedia article cites you when your own pages state a fact cleanly enough to be referenced and reached. A retrieval-grounded engine quotes you on a narrow query when the relevant passage is isolated, reachable, and unambiguous, not buried three menus deep on a page no link path leads to.
The oligarchy is real, and you do not get to vote your way into it. What you control is whether the few sources that do get cited can find a clean, reachable claim on your site to point at, and whether a retrieval engine that bothers to look can reach the same. That is a property of your link structure, and it is the same question the first brief asked: what your site looks like to something that can only follow links.
The web’s citation layer is collapsing onto a few hubs. You will not be one of them. The work is making sure that when the hubs and the retrievers look, there is something clean to cite.
Sources:
Most-cited domains and per-engine breakdown (Reddit, YouTube, LinkedIn, Wikipedia, Forbes; ChatGPT / Perplexity / Google AI Mode / Gemini snapshots; ~30 million cited sources across five engines): Peec AI analysis, reported by Search Engine Land (March 2026). Top-15 concentration (68% of citation share), the 680-million-citation / six-study synthesis, the PageRank comparison, the Claude prestige-press tilt, and the ChatGPT Reddit-share swing (~60% to 10% in two weeks, via SEMrush): 5W Public Relations, AI Platform Citation Source Index 2026, reported here as a vendor synthesis without a single disclosed methodology. Owned-media citation share (under 1.5%): SolCrys, 2026, one category snapshot. Link-reachability framing: Axiom Graph, Structural Signals. All third-party figures are reported as claimed by their sources.