How to Get Cited by ChatGPT: 4 Checks With a Test Each
How to get cited by ChatGPT: test your page against four checks, from OAI-SearchBot access to a quotable passage, and fix the one that fails first.
By Furkan AktaşPublished Updated

On this page
If your page ranks somewhere but never shows up in ChatGPT's sources, one of four checks is failing, and you can find which one in an afternoon.
Here is how to get cited by ChatGPT: a page has to be fetchable by OpenAI's crawler, rank for the question and its sub-questions, carry a title and URL that earn the open, and hold a passage worth quoting. Fix the earliest failure first.
In this guide:
- The four checks, in order, with the failure signature and the fastest test for each
- Two checks the usual advice skips: the firewall's IP rule and the title-and-URL step before ChatGPT reads your page
- Which popular tactics have evidence behind them, and which rest on thinner data than the advice suggests
- Why "retrieved to cited" is 15%, 50% or 75% depending on the study, and how long a fix takes to show
How to get cited by ChatGPT: four checks, in order
To get cited by ChatGPT, a page has to pass four checks in sequence: a crawler fetch, a search ranking, a title and URL that earn the open, and a quotable passage.
How does ChatGPT cite sources, and how does ChatGPT choose sources? A fork comes first. ChatGPT decides for itself when to search: OpenAI says it may search the web automatically when a question would benefit from current information. An answer built from training data carries no citation, so only a web search answer can name your page.
When it works with a search provider, OpenAI says ChatGPT typically rewrites your question into one or more targeted queries and sends those to the provider. That rewrite matters for check 2, because your page competes for the rewritten questions, as well as the one the user typed.
Each row below gives a failure signature (what you see when the page fails that check) and the fastest test:
The common mistake is polishing passages while check 1 or 2 is failing. A fix to a later check can't repair an earlier one.
The order also follows the research. A July 2026 survey of GEO studies found that the widely cited Princeton gains are conditional on a source already being present in a fixed context, and that they establish neither organic discoverability nor traffic. Passage tactics only matter once the first three checks pass.
Passing all four still guarantees nothing. OpenAI says search results and citations can be incomplete, outdated, or incorrect.
To read one page's result, ask ChatGPT the page's sub-questions and see which URLs it names, or watch for utm_source=chatgpt.com in referrals. The method for a full prompt panel is in our guide on how to track ChatGPT mentions. Brand-level strategy, such as which off-site mentions matter, lives in ChatGPT SEO.
Check 1: can OAI-SearchBot fetch the page?
A page passes this check when robots.txt allows OAI-SearchBot and the host or CDN lets OpenAI's published IP addresses through.
OpenAI's wording has two conditions: allow OAI-SearchBot to crawl the site, and confirm that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses. The first is easy to check. The second is where pages quietly fail.
If you've allowed the bot in robots.txt and still see nothing, check these failure points:
- A robots.txt rule catches the crawler. A Disallow or wildcard group that applies to OAI-SearchBot blocks it, and OpenAI says sites opted out of it won't be shown in ChatGPT search answers, though they can still appear as navigational links.
- The page needs JavaScript to render. We migrated our own React single-page app to prerendered static HTML for 20 marketing pages because AI crawlers do not execute JavaScript.
- The edit is under a day old. OpenAI says its search systems can take about 24 hours to adjust after a robots.txt update.
- The firewall or CDN rejects OpenAI's IP addresses. The robots.txt file can be perfect while a bot-protection rule drops the request before it arrives. OpenAI publishes the searchbot list at openai.com/searchbot.json, so compare it with your allowlist.
- A noindex meta tag or a Disallow sits on the page. OpenAI's publishers FAQ says that if it gets the URL of a disallowed page from a third-party search provider or by crawling other pages, it may surface just the link and page title in ChatGPT Atlas. A noindex meta tag only works if the crawler is allowed to fetch the page and read it.
OpenAI lists GPTBot as the training control and says ChatGPT-User is not used to determine whether content may appear in Search. Whether to allow GPTBot is a separate policy decision, covered in our guide on whether to allow GPTBot.
Now the test. It takes ten minutes:
- Open your robots.txt and find the group that applies to OAI-SearchBot. A specific group beats the wildcard, so a
User-agent: *block that looks harsh may not apply at all. - Send a curl request with the crawler's user-agent string, using the example in OpenAI's crawler documentation (the version number may change):
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
https://example.com/your-page
Then read the result carefully, because each outcome proves less than it looks:
So a success is weak evidence and a refusal is ambiguous. The decisive step is the IP comparison: if your firewall doesn't allow the published list, the page fails check 1 whatever curl says.
Check 2: does the page rank for the prompt and its sub-questions?
Ranking for the prompt and its sub-questions raises the odds of selection: among retrieved pages, 43.2% of those ranking #1 in Google were cited, 3.5 times the rate beyond the top 20.
Both figures come from a March 2026 analysis of 15,000 prompts and describe pages ChatGPT had already retrieved. They are not a measure of getting into the pool.
Rank is a predictor, not a pass/fail gate. In the same analysis, 55.8% of cited pages ranked in a Google top 20 for at least one query, so 44.2% did not (100 minus 55.8). Plenty of cited pages sit outside the top results for every query the study ran.
Sub-questions widen the ground a page can win on. ChatGPT generated two or more fan-out queries on 89.6% of searches, widening 15,000 prompts into 43,233 queries. And 32.9% of cited pages that appeared in any top-20 results page were found only through those fan-out queries.
So if your page doesn't rank for the headline prompt, it may still rank for one of the sub-questions ChatGPT generates.
You'll often read that Bing indexing is a prerequisite. OpenAI says ChatGPT search sometimes partners with other search providers and links Microsoft's and Shopify's privacy policies without saying which provider handles which queries, and the rank study above used Google only. A Bing check is cheap, but treating Bing as a gate is a guess.
Here's the test:
- Write the prompt a buyer would type, then a few sub-questions it would break into.
- Rank-check your page in Google and Bing for each.
- Ask ChatGPT the same questions and note which URLs it names when a competitor's page wins.
The fix is to pick the one sub-question your page answers best. If the page sits outside the top 20 for it, follow the steps in how to rank higher on Google for that page before touching anything else.
Check 3: do the title and URL earn the open?
ChatGPT picks which retrieved pages to open from each result's title, snippet and URL, according to Ahrefs, so write the title and URL for one sub-question.
Ahrefs calls this a gatekeeping layer before ChatGPT reads any page content, and it relays outside research into that step. OpenAI hasn't confirmed it. Treat the result card (title, snippet and URL) as a well-supported practitioner observation and not a documented mechanism.
What the data shows, from the March 2026 15,000-prompt analysis and Ahrefs' 1.4 million prompt study:
- Title overlap: pages with 50% or greater title-query overlap were cited at 20.1%, against 9.3% for pages with little overlap, a 2.2 times lift. The analysis doesn't say which query the overlap is measured against, so it doesn't prove sub-question matching.
- Natural language URL slugs: search results with them were cited at 89.78%, against 81.11% without, a gap of 8.7 points (89.78 minus 81.11).
- Closeness to sub-questions: Ahrefs measured a cosine similarity of 0.656 between cited titles and the best-matching fan-out query, against 0.602 for the prompt; the fan-out score takes the highest of several queries, so part of that gap is built in.
- Limits: Ahrefs used open-source embeddings to approximate ChatGPT, and some non-cited URLs were likely never opened.
Both figures are observational. They support the rule without proving it.
Here's the rule applied. A guide whose title reads "Our Approach to Crawler Settings" with the slug /blog/new-post-copy matches no question anyone types. Retitle it "Does OAI-SearchBot Need an IP Allowlist?" with the slug /blog/oai-searchbot-ip-allowlist and it matches one sub-question exactly.
The same logic explains why a deep product or topic page beats a homepage. The open decision reads each page's own title and URL, so a specific page can match one sub-question and a homepage title can't.
Check 4: does the page hold a passage worth quoting?
A page that clears the first three checks needs a self-contained answer in the first two sentences under each heading and at least one fact nobody else can state.
If your page is fetched, ranked and opened but still isn't quoted, this is the check to work on. Test it by copying the first two sentences under any heading and reading them alone. If they need the paragraph above to make sense, rewrite them.
You'll see capsule lengths prescribed to the word. Treat any exact length as a working definition, not a measured optimum, and judge a capsule by whether its two sentences stand alone.
Know the limit of the evidence first. The on-page audits behind most advice sample pages that already earned ChatGPT referrals, and the strongest causal result, the Princeton GEO gain, holds only once a source is already in context.
Take the answer capsule, a short direct summary under a heading. A Search Engine Land audit of blog posts that received ChatGPT referrals found capsules were rarely linked and read that as a drag from links. But the audit had no comparison group of uncited posts, so it never measured a base rate. A high link-free share can simply reflect how capsules are normally written.
The article itself closes by saying linking still matters and to treat the finding as a framework to fit your own linking standards. A link-free capsule costs nothing, so keep it as a default.
The table below sets each tactic against its strongest evidence:
Original data deserves the effort because it is the one tactic where the page holds something no other page can state. Owned framing, such as your own named method, helps a passage stand alone. Whether original data also helps ChatGPT choose the page is untested in these studies.
In Ahrefs' matched test, the ChatGPT result for JSON-LD schema was indistinguishable from zero. The full test is in our guide to schema for AI search.
Freshness needs a correction in both directions. Across 17 million citations, ChatGPT cited URLs 458 days newer than Google's organic results, so ChatGPT does prefer fresh content overall.
But inside one prompt's retrieval set, the older, more established pages tended to win, and the freshest tended to be discarded. The median cited age was about 500 days, with some pages over 2,700 days old.
Ahrefs adds a caution: its pool of non-cited pages is far smaller than the cited group, which limits confidence in the age gap. It also says freshness is non-negotiable for news. So update pages when facts change, and don't expect a new date alone to move a citation.
ChatGPT retrieval vs citation: which rate to believe
ChatGPT cites between 15% and 88.46% of the pages it retrieves, depending on what a study counts as retrieved, so a benchmark means nothing until you know the denominator. The next question, which sources does ChatGPT cite once it cites, has its own domain-by-domain answer.
ChatGPT retrieval vs citation rates come in four versions. Here they are side by side:
Download CSV (CC BY 4.0)
The 50% moves to about 75% once you take out the Reddit feed. Reddit has its own channel in ChatGPT's retrieval system, and Ahrefs counted over 16 million Reddit data points, about a third of the retrieved rows, 34% (16 million divided by 1.4 million prompts times 16.57 plus 16.58 URLs each). Reddit is cited at 1.93%, so it drags the average down.
To get the 75%, multiply the prompts by the cited and non-cited URLs per prompt, take out the Reddit rows and their 1.93% of citations, and divide what remains: 75% of the non-Reddit rows were cited. Reddit makes up 67.8% of the non-cited pool, so a plain cited-versus-non-cited comparison mostly compares search results with Reddit API output.
The 88.46% is the cite rate of web-search rows across 25,563,589 data points in Ahrefs' table. It is a different quantity from the share of cited URLs that come from search, so don't quote it as one.
Why plan on 15%? Because only the 15,000-prompt analysis defines cited as the URLs explicitly cited in the final answer, and its 5.5 cited pages per prompt (82,108 citations divided by 15,000 prompts) is about what an answer shows. A separate analysis of ChatGPT conversations found about 6 unique citations in each conversation that cites anything, close to the same figure.
Ahrefs counts 16.57 cited URLs per prompt, 3 times as many (16.57 divided by 5.5), more than an answer displays. That suggests its cited flag covers the whole source list, which is our inference. Ahrefs' own note says its 50% covers the full journey from retrieval to citation, with some non-cited URLs never opened.
Here's the rule: use 15% to size how many retrieved pages become citations, and 16.9% for how-to queries, the second-highest query type in the 15,000-prompt analysis. Never compare your own rate to a figure whose denominator you don't know. For tracking your own, AI search analytics covers what to measure.
How long does it take to get cited by ChatGPT?
OpenAI documents one delay, about 24 hours for a robots.txt edit, and every longer timeline in circulation is an estimate without a measurement.
Ranges from a few weeks to a few months circulate, and none states how it was measured, so treat them as untested. Instead, re-test each check on its own clock:
If you also want Perplexity, the timing and mechanics differ, and we cover them in how to get cited by Perplexity AI.
ChatGPT citations can be debugged, one gate at a time. If your page is invisible, start with the fetch test today, fix the earliest failing check, and re-test it on its own clock before you touch the next. Mission Growth's platform tracks AI citations and visibility for customers.
Frequently asked questions
Does ChatGPT always cite its sources?
No. OpenAI says responses that use web search may include citations, and ChatGPT decides itself when to search. Search placement isn't guaranteed, and a page that passes every check can still go unnamed. Answers drawn from training data carry no citations at all.
Do I need to rank on Google to be cited by ChatGPT?
Ranking raises the odds without being a requirement. Among retrieved pages, 43.2% of those ranking #1 were cited, but 44.2% of cited pages in the same study ranked in no top 20 for any query it ran. Treat rank as a strong aid, and check Bing as well.
Does ChatGPT cite Reddit?
Rarely. ChatGPT pulls a large Reddit feed, 67.8% of its non-cited pool in Ahrefs' data, and cites Reddit at 1.93%. Reddit shapes context far more than it earns links, so don't build a citation plan on it.
Will optimizing for ChatGPT hurt my Google rankings?
Nothing in this research measures an effect on Google rankings. Checks 1 to 3 are ordinary search work. The one related finding, that citation-oriented rewrites can impair retrieval, concerns generative engines, so rewrite section openings and keep the relevance text around them.
How do I check whether ChatGPT cites a specific page?
Ask the page's sub-questions in ChatGPT and read the URLs it names. Watch GA4 for utm_source=chatgpt.com in referral URLs, and look for OAI-SearchBot, not GPTBot, in server logs. The full method is in our guide to tracking ChatGPT mentions.
Cite this page
Aktaş, F. (2026, October 4). How to Get Cited by ChatGPT: 4 Checks With a Test Each. Mission Growth. https://missiongrowth.io/blog/how-to-get-cited-by-chatgpt
Download the data: 1 table as CSV (CC BY 4.0)
Figures we made for this post are free to reuse under CC BY 4.0 with credit to Mission Growth.
Get Mission Growth highlighted in your Google results.


