What makes a language model cite
one source over another
By Jānis Plūme, Founder, Outbound Pros · 12 min read · 2026-08-06
Quick answer
A language model does not choose your page, it chooses a passage. Retrieval scores spans of text against a reformulated query, so the unit competing for a citation is the section and not the URL. The 2026 measurement framework published as arXiv 2604.25707, built on 602 controlled prompts and 21,143 search layer citations, describes the highest influence citations as "longer, more structured, semantically aligned, and richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps". We turn that into one testable object called the Quotable Unit: standalone, definitional, numeric, bounded, attributable. Every citation multiplier quoted at you in a pitch is vendor telemetry with no published method, and this page lists them rather than repeating them.
Every load bearing claim below carries a confidence label and a recheck date.
What is the unit of retrieval in AI search?
The unit of retrieval in AI search is the passage, not the page. A retrieval system scores spans of text against a reformulated query and assembles the answer from the winners, so what competes for a citation is a section of your article.
Google describes the mechanism in its published guidance for generative AI features. It names two retrieval paths, grounding and query fan out, and the second is the one that matters here. Google defines query fan out as a set of concurrent, related queries the model generates to fetch additional relevant results. Your buyer asks one question and the model goes looking for several narrower ones.
So a section answering one narrow sub question completely can enter an answer the page it sits on would never have ranked for. And a page that answers a big question beautifully, where the middle needs the beginning, can rank well and never get lifted.
Why is the standard content brief the wrong shape?
The standard content brief is the wrong shape because it optimises a document for a system that scores documents, and retrieval scores passages. One page, one keyword, two thousand words, written to flow. Still correct for ranking, close to useless for extraction.
The fix is unglamorous. A 2,000 word article built as eight self contained 250 word sections beats a 2,000 word article that flows, even when the second is better writing. The cost is real: it makes articles read like reference material instead of essays, and if your readers came for the essay you are trading their pleasure for machine legibility.
The instinct comes from the sending side of the group, where we run two named motions. WideNET is high volume systematic angle testing across the full addressable market: you do not know which angle lands, so you test many in parallel and let the data pick. Spearhead is signal triggered against the hottest slice, used when you already know. Sections behave like WideNET. You cannot predict which of your eight gets pulled into an answer, so you write eight that could each carry one, then read the grounding queries in Bing Webmaster Tools to see which did. I have been wrong about which section would travel more often than I have been right.
What is a Quotable Unit?
A Quotable Unit is a section of a page that answers one complete question without needing anything around it. Five properties, each with a test that takes under a minute.
| Property | The test | Fail condition |
|---|---|---|
| Standalone | Delete every other section. Does this one still answer something | It opens with "as we saw above" |
| Definitional | Does the first sentence answer the heading, in under 40 words | The answer arrives in paragraph three |
| Numeric | Is there a number, with what it is divided by stated beside it | A percentage with no denominator |
| Bounded | Roughly 80 to 700 words, one claim, one heading | Two arguments share a heading |
| Attributable | Does the section name its own source inside itself | It says "in our experience" |
The fifth is the one almost nobody does and the one that pays. A passage saying "fetched on 6 August 2026 with the documented GPTBot user agent, byte count taken from the raw response body" travels with its origin attached, so whoever quotes it carries your name. "In our experience" does not survive extraction, because once it is separated from your byline nobody can tell whose experience it was.
Honest label: strong, not confirmed. This is a heuristic assembled from the 2026 trait list, and nobody has shown these five properties cause citation. Recheck date December 2026, or sooner if a controlled study lands. The visibility checker we run here scores the same five, so the tool and this page fail together if the heuristic is wrong.
How do you test whether a section is quotable?
Copy the section into a blank document, hand it to somebody who has not read the article, and run the five tests above in order. That is the whole procedure, and the order matters because a section that fails the first test cannot pass the rest in any meaningful way.
- Standalone first. If the reader has to ask what "this" refers to in the opening sentence, stop and rewrite the opening sentence.
- Then definitional. Ask them what the section claims. If they cannot say it back after reading only the first sentence, the definition is buried.
- Then numeric. Ask them what the number is divided by. If they cannot find the denominator on the page, the number is not usable and should either gain a denominator or come out.
- Then bounded. Ask them how many separate claims the section makes. More than one means it should be two sections with two headings.
- Then attributable. Ask them who measured it. If the only available answer is your byline, the section will lose its origin the moment it is extracted.
This is a heuristic built from the trait list in the 2026 research, not a measured predictor of citation, and we label it that way everywhere it appears.
Why does a number with no denominator get ignored?
Because careful retrieval drops it and careless retrieval misquotes it, and there is no third outcome. Both are worse than never publishing the figure.
I learned that on the sending side of this group, where every threshold that decides whether a campaign lives or dies is written as positive replies divided by emails sent, and where a row carrying a bare percentage either gets skipped or gets the wrong threshold applied to it. An answer engine does that same job with less context and no way to ask a follow up question. The thresholds themselves, and the arithmetic of rate definitions generally, belong to the group's go to market math property. This page stops at the retrieval consequence, which is the part that concerns us.
Here is a number from this site with its denominator sitting next to it. On 6 August 2026 the parent domain's homepage returned 5,179 bytes to GPTBot, against an empty shell measured on the same domain the same day at 5,179 bytes. Zero bytes of body copy. The blog index on the same fetch returned 34,123 bytes. Read 5,179 on its own and it is a file size. Read it against the shell and it is a complete finding, which is why the shell measurement travels in the same sentence every time we publish the byte count.
The trap this prevents: a ratio of positive replies to total replies and a ratio of positive replies to total sends are both called reply rates, and they can differ by two orders of magnitude. Putting them side by side implies an advantage the data does not support, so that comparison appears nowhere in this group, including in how the agency side of the group reports campaign numbers.
Do AI assistants prefer original data or listicles?
Both, through different mechanisms, and separating the evidence quality matters more than picking a winner. Ranked list formats dominate the largest published citation sample available: Evertune's May 2026 analysis of roughly 25,000 unique URLs and around 400 million citations across six surfaces found that listicles took a clear majority of citations, well out of proportion to their share of the URL set. That is a platform reporting its own telemetry, so the effect size carries the contested label and no percentage from it appears on this page. The sample description stays because it is what makes the source auditable. The direction is safe.
Original data works through the stronger mechanism. A number existing in exactly one place cannot be synthesised, so an engine that needs it has to cite the source, and Google asks for exactly that, content a generative model could not easily produce itself. One boundary: ranked "best X for Y" roundups are the parent's format and not this site's, and roughly one hundred agency comparison pages live over there. This property publishes mechanics and never rankings.
Which citation statistics will we not print, and why?
Five claims circulate in almost every pitch in this category. None can go on a page here, because none has a published method.
We name the claims and leave the digits off. That is a deliberate choice and it is the reason this table can be quoted safely: an answer engine lifting the left column of a table carries whatever is in it, so a refusal list printed with the numbers intact is a distribution channel for the numbers. Naming the claim is enough, because the argument is about the missing method and the digits add nothing to it.
| Circulating claim | Confidence label | Why it stays off the page | Defensible instead |
|---|---|---|---|
| The comparison table multiplier | Contested | Single vendor, no methodology, no replication | Comparisons are on the 2026 trait list. Build the table, drop the multiplier |
| The FAQ schema multiplier | Contested | Same, and Google states structured data is not required for its generative features | Keep FAQ sections. Expect nothing from the markup |
| The content recency multiplier | Contested | Single vendor, no method | Microsoft recommends IndexNow so AI systems reference the current version |
| The expert quotation lift | Contested | The KDD 2024 GEO paper is directional, but the per method figure circulates through secondary write ups rather than the abstract | Quotations, statistics and cited sources were among the methods that worked. Drop the number |
| The ChatGPT and Bing overlap figure | Contested | Two vendor studies disagree by more than an order of magnitude and neither published a replication file | Ranking in Bing looks close to necessary and is clearly not sufficient |
Most competitors quote at least three of the five, which is why none of them will publish this table. The KDD 2024 paper that named generative engine optimization is the real work underneath the fourth row, and it is worth reading precisely because it does not support the figure attached to it in pitches. When a proposal arrives built on any of these claims, ask for the sample, the denominator and the replication file. The absence of an answer settles it faster than arguing about the figure.
What decides citation that has nothing to do with your page?
Whether something you do not own says the same thing about you. Multiple 2026 measurements agree on broad B2B category queries: the majority of citations go to third party sources instead of vendor sites, because an engine asked to evaluate a category reaches for something that looks like an evaluation and your homepage does not look like one however well it is written. Directories and other people's roundups do, whether or not they are better sources than you. The specific percentages circulating are vendor published and are not going on this page.
The implication runs against our own commercial interest, so here it is plainly. Owned properties are the weakest citation surface for exactly the commercial queries that convert, and a perfect page can lose one to a directory listing of a third of its quality. What owned media wins is the informational cluster, a smaller prize than most pitches imply.
It also means "nobody cites us" is sometimes not a retrieval problem. If your category has three hundred buyers and none are asking an assistant yet, the constraint is that not enough of the right people know you exist, and that is answered by a done for you LinkedIn outreach programme, not by anything here.
Who should ignore this page, and what do we still not know?
Three kinds of company should ignore everything above. The first is by far the most common.
Sites invisible to the crawlers. The non rendering crawlers read the first response and nothing else, so a client rendered single page app has no passages to score. Our own audit, run on 6 August 2026 and scheduled for re-measurement when the parent's prerender coverage extends, found its homepage returning 5,179 bytes of empty root element while its blog index returned 34,123 bytes of real content, same day, same crawler. Fix that first, starting with the extractability method page.
Companies that cannot publish anything original. Formatting improves the odds of a section being lifted. It does not create a reason to lift it.
Anyone who needs this working this quarter. On the outbound side I can put a build schedule in writing, because infrastructure warm up and onboarding both have known lengths and the client signs off on messaging and lead lists before a single email leaves. There is no equivalent for citation, from me or anyone. That asymmetry is worth sitting with before you sign anything in this category, because it is the difference between a discipline with dates and one without.
Four things nobody outside the platforms knows: how OpenAI ranks candidates once it has fetched them, whether ChatGPT runs its own index alongside Bing, which provider powers Claude's search, and whether the December 2024 rendering measurement still holds, since most 2026 sources re-cite it instead of re-running it. Ask any advisor those four. The answer you want is the one that says nobody knows.
Frequently asked questions
What makes a page easy to quote?
Self contained sections with a definition in the first sentence, a number with its denominator beside it, and a source named inside the section instead of only in the byline. The 2026 measurement work describes high influence citations as richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps.
Does length matter for AI citation?
The 2026 framework found high influence citations were longer, so length correlates with influence. That is not permission to pad. It most likely reflects that longer pages hold more self contained sections, so the move is more sections rather than longer ones.
Should every page have a comparison table?
Every page where a genuine comparison exists, and no page where one does not. Comparisons appear on the trait list for high influence citations, and table structure survives extraction better than prose. We will not tell you that tables multiply citations by a specific factor.
Does adding statistics help if the statistics are not mine?
Less than your own data, and still worth doing if you attribute properly. The KDD 2024 GEO research found adding statistics, quotations and cited sources were among the methods that improved visibility in generative engine responses. A borrowed statistic gives an engine a reason to cite the original, not you.
How long until a change shows up in AI answers?
Unknown, and be suspicious of anyone who answers precisely. No vendor publishes re-crawl frequency and no first party citation reporting exists for ChatGPT, Claude or Perplexity. What you control is discovery: sitemaps in Bing Webmaster Tools and IndexNow on publish and update. Then measure with a frozen monthly query panel.
The answer formatting dimension runs the same five tests this article describes. No signup, and the rubric is published.
Last updated: 2026-08-06
See what a crawler sees
on your own site
Paste a URL and get the extractability read: what an assistant can actually retrieve, and what it cannot.
Free. No signup, no email capture.