Advice on getting cited in ChatGPT usually comes from reading its answers and guessing at the machinery behind them: write quality content, publish listicles, be active on Reddit. Suganthan Mohanadasan took a different route. He opened the browser’s network panel and read the traffic ChatGPT sends itself. What comes back is a more mechanical picture of how the model chooses which pages to fetch, which to cite, and which to ignore. This article retells his capture, adds the one piece of outside data that backs it, and says what it supports and what it does not.

Update, July 2026: the pipeline field described below later showed a fifth value, bing, on some accounts. That follow-up is covered separately in ChatGPT and the Bing result source.

First, a word on the method and its limits

This is one person’s read of roughly 1,240 source records, captured over a few days on a single Pro account, with queries skewed toward SaaS and tech. Suganthan splits his findings into two confidence levels, and we keep that split.

The structural facts are high confidence: they come straight off the wire, the same way every time. The internal field names, the pipeline labels, the fact that some questions never trigger a web search at all. The frequencies are directional only: the percentages and rankings come from a small, tech-heavy sample, and health or fashion queries would look different. Read the mechanisms as the finding and the exact numbers as a snapshot.

Not every question even touches the web

Before ChatGPT searches anything, it files your question into a bucket. Suganthan saw six of these intent labels, including one that settles a lot of arguments about AI visibility: the “text” bucket answers from the model’s training memory and never searches the web at all.

That has a blunt consequence. Some questions you simply cannot win with content, because no page is ever fetched to answer them. In his test, three of ten deliberately current, high-stakes questions were answered with no web search. And the bucket is decided by how you phrase the question, not the topic. “Best coffee near me” runs the local pipeline. “Best 4K TVs to buy” triggers shopping. “Best 4K TVs with reviews” stays in normal search. Same subject, three different machines, three different sets of winners.

So before you worry about ranking in an AI answer, work out whether the questions you care about reach the web at all. Some never will.

Four pipelines decide who gets fetched

For the questions that do search, every web result carries an internal label saying which pipeline fetched it. Suganthan found four values, and others in the industry surfaced the same field independently. In plain terms:

  • serp is the open-web baseline, mostly news.
  • labrador looks like a licensed tier for established publishers (think Reuters, the Guardian, the WSJ, Wikipedia, arXiv), pulled in as near-complete article extracts rather than short snippets.
  • bright and oxylabs are the names of commercial scraping firms, Bright Data and Oxylabs, and they do the heavy lifting on the open web. In his sample, Bright Data dominated shopping, finance and weather, while Oxylabs leaned toward regional and local press.

A single weather query, for example, split its sources across pipelines: the national forecasters came through one scraper, the regional papers through another. You do not control which pipeline picks you up, but the pattern is a reminder that ChatGPT is assembling an answer from several fetching systems at once, not searching one index like Google.

Fetched, cited, and mentioned are three different things

Most visibility advice blurs this distinction. Being read by the model, being credited under a specific sentence, and having your brand name surface in the answer are three separate outcomes, and only one of them is the clickable footnote people call a citation. Each has its own win condition:

  • Fetched means the model pulled your page into its working context. You never see this, and it earns you nothing on its own.
  • Cited means your page is credited as the source for a specific sentence, the clickable footnote.
  • Mentioned means your brand name appears in the answer, often as a chip, without being the source of any claim.

The gap between fetched and cited is where a lot of effort is lost. In Suganthan’s sample, Reddit was fetched 278 times and cited 11 times. YouTube was fetched 201 times and cited zero. The mechanical reason is simple: citations bind to text the model actually read, and a YouTube page returns metadata, not a transcript, while a Reddit thread is all readable text. This is not a small-sample fluke. Ahrefs, studying 1.4 million prompts, found Reddit cited only about 1.93% of the time it appeared, and that roughly two thirds of all fetched-but-uncited URLs were Reddit. The platform is used constantly to understand a topic and almost never credited for it.

A console table of the most-cited domains in the sample, led by reddit.com, rtings.com, zoho.com, semrush.com and techradar.com
The most-cited domains in Suganthan's sample: Reddit narrowly on top, then review hubs and vendor pages. Screenshot from Suganthan Mohanadasan's analysis, suganthan.com. The data is his, from a small tech-skewed sample, not ours.

One more mechanic hides in that list: results deduplicate by domain. Twenty thin pages from one site collapse into a single citation slot. Publishing more of the same page does not buy you more presence.

The same split between being cited and being recommended shows up in Google’s AI Overviews. Why ‘We’re the Best’ Backfires in AI Search covers that measurement.

One question becomes dozens of searches

When ChatGPT uses its “thinking” mode on a comparison question, it does not run your query. It runs many. Suganthan watched single comparison tasks fan out into fifteen to forty sub-queries, and because those sub-queries are logged, you can read exactly what the model asked.

The behavior is strikingly literal. It fires site: probes straight at vendors’ pricing pages. It guesses a price, then searches to confirm the guess. It goes off-script and pulls in competitors you never mentioned, then hunts for their pricing too. When it reads a page, it is grepping for dollar signs and specific words like “Agency” or a plan name.

A list of the many sub-queries ChatGPT generated from a single comparison question, each targeting a different vendor's pricing
One comparison question, fanned out into many machine-written sub-queries. Screenshot from Suganthan Mohanadasan's analysis, suganthan.com. The queries are his capture, not ours.

The lesson for visibility is uncomfortable: you are not competing for the question a user typed. You are competing for a swarm of rewritten sub-queries you never see, many of them aimed at pages you may not have optimized, like your own pricing page.

It reads your page for facts, everyone else’s for opinion

When Suganthan looked at the thinking model’s own saved reasoning, the strategy was written out in plain language. For hard facts like pricing and specs, it prefers the official source and says so, choosing a current pricing page over an older third-party mention. For judgment calls like “which tool is best,” it sources the verdict to third parties: review hubs, comparison sites, communities.

His own summary is the one sentence to keep from the whole capture: ChatGPT “reads your own page for the facts, if it can parse them, and everyone else’s for the opinion.” Two different jobs, two different sources. You own the facts about yourself. You do not own the verdict, and the verdict is earned elsewhere. The other articles in this series on AI search build on that split and link back here rather than restate it.

The JavaScript wall

This finding matters most for how pages are built. When the model went to pull pricing off tools like Profound and Peec, it could not. It noted, in its own reasoning, that the prices were not in the result because the page loaded them with JavaScript. Unable to parse the official number, it fell back to quoting a third party instead, deciding to use figures from a review site because the official page was too hard to read.

For facts, this is the whole game. If your key numbers, your prices, your specs, your availability, are rendered by JavaScript after the page loads, the model may fetch your page, fail to read the fact, and hand the citation to whoever wrote about you in plain text. You can be the definitive source for a fact about your own business and still lose the citation to a comparison site, purely because of how your page is built.

What to change on your own site

The capture supports two kinds of action: make your facts readable, and do not expect volume to help.

  • Put your facts in plain, server-rendered HTML. Prices, specs, locations, hours, plan names. If a fact matters, it should be in the page text, not painted in by JavaScript after load. This is the fix the capture supports most directly.
  • Keep one strong page per fact. Deduplication by domain means twenty near-duplicate pages collapse into one citation slot. One clear, canonical pricing or product page is enough.
  • Know which questions never search. If your buyers ask things the model answers from memory, no content wins them. Aim your effort at the questions that trigger a fetch.

The other half, the verdict, is not something a page edit fixes. It comes from reviews, community discussion and independent comparisons. That is the work behind AI visibility, and it is why a self-written “best” page does not carry over.

You can check your own, no special access needed

You do not need to be a researcher to see some of this. In ChatGPT, open your browser’s developer tools, switch to the Network tab, turn on “preserve log,” run a query, and search the responses for the field name result_source. You will see which pipeline fetched each result on your own sessions, and nothing leaves your machine. Suganthan documents the full method, and Olivier de Segonzac has published a free Chrome extension that captures the same fan-out and pipeline data and exports it to a spreadsheet. We have not run this capture on cyberlab.team, so every number above is his, not ours.

What this snapshot supports

It is one account, a few days, a tech-heavy set of queries, and a system OpenAI changes constantly. The percentages will drift, as Suganthan says himself, so lean on the structure: ChatGPT sorts the question into a bucket, fetches through pipelines you do not control, reads your page for facts only if it can parse them, and takes the verdict from other people’s pages. It does not support any claim about how many citations a given fix will win you, and it says nothing about Google’s AI features.

Sources

  • Suganthan Mohanadasan, “How ChatGPT Actually Picks Sources (I Read the Network Traffic, Not the Outputs),” 24 June 2026: suganthan.com
  • Ahrefs, “Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts)”: ahrefs.com
  • Olivier de Segonzac, free ChatGPT search fan-out Chrome extension (writeup): think.resoneo.com