Mission Growth

AI Citation Statistics 2026: Top Domains and the 68% Myth

AI citation statistics traced to primary studies: Wikipedia is 7.8% of ChatGPT citations, the 68% top-15 claim fails, Claude cites 13 sources when it cites.

By Published

Ten clear quote cards stand in front of a far larger green pile spilling from a chrome box, the long tail behind AI citation statistics
On this page

AI citation statistics are easy to find and hard to compare. A writer who wants one number for a deck usually finds three for the same domain: Wikipedia at 7.8% of ChatGPT citations, at 47.9%, and at 55%. All three are real, and they sit on different bases.

This page converts the study shares into one base and traces three circulating figures back to their primary study. A 68% top-15 share has no published base or computation. A 44.2% intro share is credited to the wrong study and scoped to ChatGPT. A Reddit 60% is a share of responses.

It is a child of our ai seo statistics hub, so these AI citation stats cover bases, corrections, sources per answer, ghost citations and accuracy, plus the AI source citation preferences each engine shows. Wondering what are AI citations? They are the sources an AI answer surfaces, and the FAQ defines them. The ranking-impact studies stay on the hub.

Key AI citation statistics at a glance

AI citation statistics describe different bases, so every number below carries its base, sample and date; we last verified them on 5 October 2026. These ten are safe to lift whole:

  • Wikipedia is 7.8% of all ChatGPT citations in Profound's dataset of 680 million citations (August 2024 to June 2025).
  • ChatGPT's ten most-cited sources hold about 16.3% of its citations, and Perplexity's hold about 14.1%, by our conversion of Profound's published figures.
  • Earned media is 84% of AI citations in Muck Rack's May 2026 edition, which analysed more than 25 million links from ChatGPT, Claude and Gemini.
  • Yext's October 2025 study of 6.8 million citations finds 86% from sources brands control.
  • In a sample of 18,012 ChatGPT citations (Growth Memo, February 2026), 44.2% came from the first 30% of the text.
  • Claude cites in 55% of responses and averages 13 sources when it does, so about 7.2 per response (Muck Rack, May 2026).
  • ChatGPT cited Reddit in close to 60% of prompt responses early in Semrush's 13-week window (14 July to 12 October 2025) and around 10% by mid-September.
  • In Semrush's June 2026 sample of 3,981 domain appearances, 61.7% were ghost citations: a source link with no brand name in the answer.
  • Between 50% and 90% of LLM responses were not fully supported by the sources they cite in a Nature Communications study (April 2025) of seven models.
  • Ahrefs found that 37.9% of URLs cited in AI Overviews also appeared within the first 10 blocks of the results page (March 2026).

What an AI citation share counts

An AI citation share is a fraction whose base differs by study (all citations, a top-10 list, prompt responses, source type or text position), and the base decides which question the number answers.

Profound's August 2024 to June 2025 dataset shows the problem on one domain: Wikipedia is 7.8% of ChatGPT's total citations and 47.9% of its top 10 most-cited sources. Semrush counts a third thing, a share of responses: the prompt responses that cite a domain, where Wikipedia sat at roughly 55% before dropping below 20%.

BaseWhat the number divides byExampleStudy
All citationsEvery citation in the datasetWikipedia 7.8% of ChatGPT citationsProfound, August 2024 to June 2025
Top-10 listThe ten most-cited sources of one engineWikipedia 47.9% of ChatGPT's top 10Profound, August 2024 to June 2025
Prompt responsesResponses that cite the domain at least onceWikipedia roughly 55% of ChatGPT responsesSemrush, July to October 2025
Source typeCitations grouped by owner or formatEarned media 84%Muck Rack, May 2026
Text positionWhere in the page the cited passage sitsFirst 30% of text 44.2%Growth Memo, February 2026
Results-page overlapCited URLs that also rank in the first 10 blocks37.9%Ahrefs, March 2026

Download CSV (CC BY 4.0)

You can recover an unknown base from two published numbers for the same domain. Divide the domain's share of all citations by its share of the top 10 and the result is the top-10 total.

For ChatGPT, 7.8% divided by Wikipedia's 47.9% gives 16.3%. For Perplexity, Reddit's 6.6% of all citations divided by its 46.7% of Perplexity's top 10 gives 14.1%. The 46.7% is Reddit's value in Profound's own Perplexity top-10 table, which 5W's May 2026 report repeats.

That conversion turns any quoted top-10 share into a share of all citations, provided the study prints the domain both ways.

Overlap is one more base, and it moves with the sample. Ahrefs matched 4 million AI Overview URLs from 863,000 keyword SERPs and found 37.9% of cited URLs inside the first 10 blocks. Our own capture reads differently.

Across 58 US Google results pages we captured between 22 September and 1 October 2026 (searches for the SEO, GEO and AI-marketing topics we research, not a random sample of US queries), 42 showed an AI Overview, which is 72.4%; the median Overview cited 7 sources, and 50.4% of the 268 cited sources were domains that also ranked in that page's organic results, or 40.7% when matched by exact URL.

For more AI overview citation statistics, our google ai overviews statistics page holds the full table of overlap readings.

Quote a share only with its base. The pipeline stage matters too: a page can be retrieved, cited, named in the answer or supported by the source, and each is a separate count (the last section returns to this).

Which domains do ChatGPT, Perplexity and Google AI Overviews cite most?

Wikipedia leads ChatGPT with 7.8% of all its citations and Reddit leads Perplexity with 6.6% of all its citations, but Ahrefs now finds YouTube the most cited domain in Google AI Overviews, so the leader depends on window and base.

Among the most cited domains in AI search, the leaders hold single-digit shares of all citations in Profound's data. Which sources does ChatGPT cite most? Wikipedia, in both Profound's and Muck Rack's data. Our chatgpt statistics page covers its wider behaviour, and perplexity ai statistics does the same for Perplexity.

The table lists the top domains cited by ChatGPT and its peers with the base each number sits on, because AI platform citation patterns only compare on a shared base.

EngineLeading domainNumberBaseStudy and window
ChatGPTWikipedia7.8%All citationsProfound, August 2024 to June 2025
PerplexityReddit6.6%All citationsProfound, August 2024 to June 2025
Google AI OverviewsReddit2.2%All citationsProfound, August 2024 to June 2025
Google AI OverviewsYouTubeMost cited domain, no total share publishedAhrefs Brand RadarAhrefs, March 2026
Google AI ModeLinkedInNearly 15%Prompt responsesSemrush, 14 July to 12 October 2025
ChatGPT, Claude, GeminiWikipedia, PubMed Central, RedditLeading domain onlyTop cited domain per engineMuck Rack, May 2026

Download CSV (CC BY 4.0)

The top 10 holds 16.3% of ChatGPT citations and 14.1% of Perplexity's, from the conversion above. That leaves 83.7% of ChatGPT's citations and 85.9% of Perplexity's in a long tail outside the ten most-cited domains.

Google AI Overviews stays unconverted, because Profound prints its share-of-top-10 value only in a table. Muck Rack's leaders are named without a share, so that row cannot sit on the same base as the others.

AI citation statistics for two engines: the top ten sources hold 16.3% of ChatGPT citations and 14.1% of Perplexity citations, leaving the rest in a long tail.
From August 2024 to June 2025, ten sources held 16.3% of ChatGPT citations and 14.1% of Perplexity’s, with the rest in a long tail, by Profound’s figures.

Concentration is real in direction. Wikipedia and Reddit lead, and a handful of domains recur across studies. At the level of all citations it is a few points per domain, which is a much flatter picture than the one the 68% figure paints.

Top domains cited by LLMs

The top domains cited by LLMs are Wikipedia for ChatGPT, PubMed Central for Claude and Reddit for Gemini in Muck Rack's May 2026 data.

Profound's August 2024 to June 2025 data has Reddit leading both Perplexity and Google AI Overviews, so the answer changes with the study.

Semrush AI sources

Semrush's AI sources data counts the share of responses that cite each domain: across 230,000 prompts from 14 July to 12 October 2025, LinkedIn appeared in nearly 15% of Google AI Mode responses.

Those are response shares, so they sit on a different base from the all-citations shares in the table above.

Do 15 domains really take 68% of AI citations?

No published base or computation supports a 68% top-15 share: it comes from a rank-weighted index of 50 domains, while Profound's published figures imply a whole ChatGPT top 10 holds 16.3% of its citations and Perplexity's 14.1%.

The claim reads, in the index's own words:

The top 15 domains capture 68% of consolidated AI citation share. Reddit alone accounts for roughly 40%.

Everything-PR published that index in May 2026 (its page now reads updated August 2026), and 5W's release repeats it as "Reddit is the #1 source across every major AI engine, cited at roughly 40% frequency across LLMs." Four checks show what the figures can and cannot be.

  1. Method. The index averages each domain's rank position, using every study's citation count as the weight, into a ranked list of 50 domains. A rank-weighted index orders domains. It does not divide one domain's citations by all citations, so it has no base from which a share could come.
  2. Dataset identity. The index states more than 680 million citations drawn from six studies between August 2024 and April 2026. Profound's single dataset is 680 million citations, from August 2024 to June 2025. The index's headline size therefore matches one study's dataset, so the figure describes inputs and does not support a share.
  3. Per-engine arithmetic. Profound's published figures imply that ChatGPT's whole top 10 holds 16.3% of its citations and Perplexity's 14.1%. Fifteen domains cannot hold 68% of a pipeline when ten hold under a fifth on both engines we could convert.
  4. Reddit across four studies. Reddit is 2.2% of Google AI Overview citations and 6.6% of Perplexity's in Profound's data, 2% of citations in Yext's study, and close to 60% of ChatGPT responses early in Semrush's 13-week window. That last figure is a share of responses, and it fell by about 50 points to around 10% by mid-September. Everything-PR itself says the 40% is the multi-engine aggregate.

What to quote instead: name the engine, the base and the window, for example "Reddit is 6.6% of Perplexity citations (Profound, August 2024 to June 2025)". If you need the index, describe it as a ranking of 50 domains, which orders them and measures no share.

Are most AI citations earned media or brand-owned pages?

Earned media is 84% of AI citations in Muck Rack's May 2026 edition, which analysed more than 25 million links from ChatGPT, Claude and Gemini across 17 industries.

Yext's October 2025 release finds 86% of 6.8 million citations from ChatGPT, Gemini and Perplexity come from sources brands control. The query designs, class schemes, windows and engine sets differ, so the two numbers describe different populations.

What is AI reading?

Muck Rack's "What is AI reading" study answers its own question with earned media: 84% of AI citations, with journalism alone 27% of cited sources and paid or advertorial content 0.3%.

Yext's 86% brand-controlled figure and why the studies differ

Yext reports that 86% of citations come from sources brands already control, such as websites and listings, and that Reddit made up 2% of citations in its study.

StudyHeadlineSampleQuery designClass schemeEngines
Muck Rack, May 202684% earned media; 27% journalism; 0.3% paidMore than 25 million linksPrompts across 17 industriesEarned versus paidChatGPT, Claude, Gemini
Yext, October 202586% from sources brands control6.8 million citationsQueries set in a specific location with a specific intentBrand-controlled sources versus the restChatGPT, Gemini, Perplexity
Otterly, January to February 202652.5% (labelled two ways)1+ million citationsNot statedBrands, news, communityChatGPT, Perplexity, Google AI Overviews

Download CSV (CC BY 4.0)

Both headline numbers can be true because "earned" and "brand-controlled" slice the population differently, and the query sets differ: Yext asks location- and intent-specific questions, which pull in listings and local pages.

Muck Rack's number is stable. Earned media has ranged from 82% to 89% across three editions, and journalism between 25% and 27%. Subtract 84% and 0.3% from the whole and everything neither earned nor paid is at most 15.7%.

Otterly's January to February 2026 report is unusable until its labels agree. Its summary says community platforms capture 52.5% of citations versus 47.5% for brand domains.

Its body says brands represent 52.5% of all citations, with news sites at 20.3% and community forums at 5.9% sharing the other 47.5%. The same split is assigned to community in one place and brands in another. The named categories add to 26.2%, which leaves 21.3 points of the 47.5% in unnamed "other sources".

Pick the study whose query design matches your queries: Yext for local and intent-led queries, Muck Rack for category-level prompts across industries. To find which prompts you lose to competitors, the method is in our ai citation gap analysis guide.

How many sources does each AI engine cite per answer?

ChatGPT cites sources in 96% of responses and averages 5, Gemini cites in 82% and averages 8, and Claude cites in 55% but averages 13 when it does. Muck Rack's May 2026 edition publishes all three pairs, and the Claude average carries a condition that the headline drops: it applies only to responses in which Claude cites.

EngineCites sources inPublished averagePer response, all responses
ChatGPT96% of responses5 citations per responseNot converted
Gemini82% of responses8Not converted
Claude55% of responses13 sources when it citesAbout 7.2

Download CSV (CC BY 4.0)

"Claude averages 13 sources per response" therefore overstates it. Multiply the cite rate by the average (0.55 x 13) and Claude gives about 7.2 per response across all responses.

We keep ChatGPT's 5 and Gemini's 8 as published, because the source does not state them as conditional and we do not multiply them. For usage context on the engine, see our claude ai statistics page.

Where on the page do AI citations come from?

In a sample of 18,012 ChatGPT citations, 44.2% came from the first 30% of text, 31.1% from the 30% to 70% band and 24.7% from the last 30%. Kevin Indig's Growth Memo published the study in February 2026.

It started from a universe of 1.2 million search results and AI-generated answers, drew on Gauge data of about 3 million ChatGPT answers and 30 million citations, and kept only the passage matches at or above a cosine similarity of 0.55.

The three bands cover different amounts of text, so raw shares mislead. Per 10% of text, the first 30% earns 14.7% of citations (44.2 / 3), the middle 40% earns 7.8% (31.1 / 4) and the last 30% earns 8.2% (24.7 / 3). The intro is 1.9 times denser than the middle band.

Band of the textShare of citationsPer 10% of text
First 30%44.2%14.7%
30% to 70%31.1%7.8%
Last 30%24.7%8.2%

Download CSV (CC BY 4.0)

Bar chart of ChatGPT citations per 10% of page text: 14.7% in the first 30%, 7.8% in the middle 40% and 8.2% in the last 30% of the text.
Per 10% of text, the first 30% of a page earns 14.7% of ChatGPT citations, the middle 40% earns 7.8% and the last 30% earns 8.2%.

The figure is a share of ChatGPT citations, so it says nothing about other engines. One 2026 roundup credits it to SparkToro (January 2026), and the origin is Growth Memo's February 2026 analysis of Gauge data.

Do not mix it up with the other 44.2% in circulation on our get cited by chatgpt page, which counts cited pages that ranked in no top 20. For a claim about page structure, write "ChatGPT drew 44.2% of 18,012 sampled citations from the first 30% of the text".

Does a citation mean the brand gets named?

No: in Semrush's June 2026 sample of 3,981 domain appearances, 61.7% were ghost citations, a source link where the brand name never appeared in the answer. Another 13.2% were both cited and mentioned, and 25.1% were brand mentions without a citation. The sample covers 115 prompts run across 14 countries and four AI search engines, which is small, so read the engine contrasts as indications.

Stacked bar of 3,981 domain appearances: 61.7% ghost citations, 13.2% cited and mentioned, and 25.1% mentioned without a citation.
Of 3,981 domain appearances in Semrush’s sample, 61.7% were ghost citations, 13.2% were cited and mentioned and 25.1% were mentioned without a link.

Engines differ in which side they favour. ChatGPT cites brands 87% of the time but mentions them in only 20.7% of answers, so it cites 4.2 times as often as it names. Gemini does the opposite: it mentions brands in 83.7% of appearances but generates a citation link 21.4% of the time, so it names 3.9 times as often as it cites.

One report points the other way. 5W's May 2026 report puts ChatGPT's brand mentions at about 3.2 times its linked citations. Semrush's per-appearance counts for ChatGPT run the opposite direction, so the two cannot both stand on one base.

Use Semrush when you need a sample you can state, and use the 3.2 times figure only with 5W's name on it.

A citation therefore does not carry brand recognition. A tracking setup that counts links alone will miss the 25.1% of appearances where the answer names you without linking, and one that counts mentions alone will miss the cited pages.

Mission Growth's platform tracks AI citations and visibility for customers.

How often do AI citations support the answer?

Cited sources frequently fail to support the answer: between 50% and 90% of LLM responses were not fully supported by the sources they cite in a Nature Communications study of seven models. A citation shows retrieval, and whether the source backs the sentence beside it is a separate question. Four studies measure it, on different tasks, so we list them side by side and do not average them.

StudyTaskSampleResult
Nature Communications, April 2025Medical questions, seven LLMs800 questions, 58,000 statement-source pairsBetween 50% and 90% of responses not fully supported
Nature Communications, April 2025 (doctor check)GPT-4o with retrievalExpert review of responses40.4% of responses fully supported
Tow Center, March 2025Identify article, publisher and URL from an excerpt that Google returns in its first three results1,600 queries, 8 chatbots, 20 publishersMore than 60 percent incorrect; Perplexity 37 percent; Grok 3 94 percent
Liu et al., 2023Generative search engines' sentences and citationsSentence-level human evaluation51.5% of sentences fully supported; 74.5% of citations support their sentence

Download CSV (CC BY 4.0)

The Nature paper adds that even for GPT-4o with web search, approximately 30% of individual statements are unsupported and nearly half of its responses are not fully supported.

The limits matter. The Nature data are medical queries, the Tow Center task is a source-finding test, and the 2023 figures predate current engines. None of them is an average of AI accuracy, and the Tow Center's incorrect answers are a different measure from the Nature "not fully supported" share. The AI Overview accuracy figure lives on the sibling page linked above.

Separate "cited" from "supported" in any tracking you run.

Which AI citation statistic should you quote?

Quote the statistic whose base and stage match your claim: a share of all citations for market size, a share of responses for reach, a rate per prompt for visibility, a support rate for trust. A 2026 survey of 45 GEO studies by Martinez models retrieval, citation, prominence and absorption as separate stages of a pipeline.

So a retrieved page, a cited page, a named brand and a supported statement are four counts that do not substitute for each other. Where a study counts cited, mentioned, or supported appearances, it measures a different stage of the pipeline. The survey also notes that commercial audits find low source overlap and run-to-run variability.

Your claimQuoteBaseStage of the pipelineDo not quote
A domain is a big source for an engineWikipedia 7.8% of ChatGPT citations (Profound)All citationsCitedWikipedia 47.9%, which is a share of the top 10
A domain reaches many answersReddit close to 60% of ChatGPT responses, early in the 13-week window (Semrush)Prompt responsesCitedAny response share as a share of citations
Earned media dominates84% earned (Muck Rack, May 2026) or 86% brand-controlled (Yext)Citations by source typeCitedOtterly's 52.5%, until its labels agree
Intro text gets cited44.2% of 18,012 ChatGPT citations in the first 30% (Growth Memo)Text positionCitedThe figure for any other engine
Citation brings brand recognition61.7% ghost citations (Semrush)Domain appearancesCited versus namedThe 3.2 times claim without 5W's name on it
Citations are trustworthyBetween 50% and 90% not fully supported (Nature)ResponsesSupportedA single accuracy average across the four studies
Top domains hold most citations16.3% of ChatGPT citations in its top 10All citationsCitedA 68% top-15 share

Download CSV (CC BY 4.0)

Crawl access is the precondition for any of these counts, because a page an AI crawler cannot fetch cannot be cited; our AI crawler statistics page covers that side.

How these numbers were checked

Each figure here comes from the study that published it, with one exception: the 3.2 times mention figure is 5W's own. We record the window, unit and sample for every study, and a share carries the base it divides by.

Muck Rack's figures come from an archived copy of the publisher's page.

The derived values (16.3%, 14.1%, 7.2, 1.9 times and the others) show their formula where they first appear. This page was last verified on 5 October 2026.

An AI citation share is not one metric. The next time you copy a number into a deck, write its base beside it, and if you can't recover the base from the study, leave the number out.

Frequently asked questions

Is being cited the same as being mentioned?

No. In Semrush's June 2026 sample, 61.7% of 3,981 domain appearances were ghost citations with a source link and no brand name in the answer, 13.2% were both cited and mentioned, and 25.1% were mentions with no citation. The two signals overlap only a little.

Are AI citations accurate?

Often not fully. A Nature Communications study found between 50% and 90% of LLM responses were not fully supported by the sources they cite, on medical questions. The Tow Center's source-finding test returned incorrect answers to more than 60 percent of queries. The tasks differ, so the numbers are not interchangeable.

How many sources does ChatGPT cite per answer?

ChatGPT cites sources in 96% of responses and averages five citations per response, as Muck Rack published them in May 2026. The source does not state the average as conditional, so we leave it unconverted. Gemini averages eight and Claude averages 13 only when it cites.

What percentage of AI citations come from Reddit?

It depends on the base. Reddit is 2.2% of Google AI Overview citations and 6.6% of Perplexity's in Profound's data, and 2% in Yext's study. Semrush found close to 60% of ChatGPT responses cited Reddit early in its 13-week window, which is a share of responses.

Which websites does ChatGPT cite most?

Wikipedia leads in two studies: it is 7.8% of ChatGPT citations in Profound's data, and Muck Rack names it ChatGPT's top cited domain. Semrush found it in roughly 55% of ChatGPT responses before a fall to under 20%. Always state the base.

What are AI citations?

AI citations are the sources surfaced in an AI-generated answer, which is how Yext defines them in its study of 6.8 million of them. A citation is a link or source reference, and a mention is the brand's name in the text. The two can appear together or apart.

Figures we made for this post are free to reuse under CC BY 4.0 with credit to Mission Growth.

Get Mission Growth highlighted in your Google results.

Related

Next step

Put these playbooks to work

Start with a free audit. See where the lift is before you commit.

How it works

  1. 01

    30-minute audit call

    We map your funnel against your goal and pull live data from your channels.

  2. 02

    Lift estimate

    You get a written estimate of where the lift is, with a 30-day plan to capture it.

  3. 03

    You decide

    Run it with us, run it in-house, or shelve it. No commitment from the audit.