CoreCited

How an AI assistant decides which sources to cite

What is actually known about how AI answers pick their sources, what is inferred from observation, and what is simply marketing.

10 min read

Nobody outside the engines knows how citation selection works. What exists is a small amount of documented behaviour, a larger amount inferred from watching outputs, and a great deal of confident writing with nothing behind it. This post sorts the three.

The mechanism, as far as it is public

A generative answer is produced in two steps. The engine retrieves documents — from a search index, from a live fetch, or both — and then a model writes prose grounded in them, attaching links to the sources it drew on.

Featured snippet
One page
quoted verbatim

Extracted, not written. One page wins and you can read the winning passage on it.

AI Overview
written

Composed from several. Nothing on any page tells you why yours was chosen.

Both halves matter and they fail differently. If retrieval never surfaces your page, nothing else you do is relevant. If retrieval surfaces it but the model cannot find a clean statement to ground a sentence in, it uses somebody else's page instead — and you will never know it was considered.

What is actually known

Retrieval overlaps with ranking, partially

Cited sources correlate strongly with pages that rank well for the query. That is the most robust observation in this whole area, and it is why conventional SEO fundamentals still matter. But the overlap is partial in both directions: pages well down the results get cited, and pages ranking first often are not.

Queries get decomposed

A multi-part question is answered from several sub-searches rather than one. This is why a page that answers one clause exceptionally well can be cited for a query it would never rank for as typed — and why targeting the exact phrasing of a question is less useful than covering the thing it is asking about.

Crawler access is a precondition

The AI user agents are separate from Googlebot. A robots.txt written before they existed does not mention them, and plenty of sites block them without knowing. We track 13 of them in the free crawler checker; it takes under a minute and occasionally reveals that a site has been invisible for months.

What is inferred, honestly labelled

These come from observing many answers rather than from anything an engine documented. They are worth acting on and should not be stated as mechanisms.

  • A self-contained answer near the top does better. The most consistent pattern anyone reports. It is also just good writing, which is a point in its favour: you are not betting on a quirk.
  • Specificity survives summarisation. Numbers, dates, named limits and constraints can be attributed to you. Adjectives cannot. “Fast and affordable” carries nothing into a generated sentence.
  • Freshness matters more on browsing engines. A page published last month can be cited within days by an engine that reads the live web, and go unmentioned for a year by one answering from training.
  • Entity clarity helps. If three pages describe your company three different ways, a model assembling a sentence about you has three weak signals instead of one strong one.

The finding brands find hardest

A large share of what gets cited is not owned by anyone in your industry. Comparison articles, roundups, review sites and forum threads make up a substantial portion of the sources in commercial answers — because a model writing a balanced recommendation needs several viewpoints, and a discussion supplies them in one document while a vendor page supplies one and reads like advertising.

So your most valuable page is frequently one you cannot edit. If the comparison article an engine reads every cycle names four competitors and not you, being added to that article beats anything you can publish on your own domain. That is the single biggest difference between this work and traditional SEO, and it is why knowing which sources get cited is more actionable than knowing your own score.

What is marketing

Three claims circulate widely and none of them has support.

  • “Citation scores.” A number claiming to predict or explain why an engine picked a source. The selection is not observable from the output, so any such score is a model of a mechanism nobody outside the engine has seen — presented as a reading of it.
  • “Submit your brand.” There is no index to join and no placement to buy. This is the most commonly sold non-existent service in the category.
  • “AI traffic analytics.” No API exposes clicks from AI answers, Search Console folds them into organic, and analytics receives them as plain Google traffic. A precise chart of AI traffic is a model wearing a measurement's clothes.

The test that sorts a vendor quickly: ask where the number comes from. A real answer names a source you could check. An unreal one describes a proprietary methodology.

What to actually do

  1. Check the crawlers can reach you. Cheapest possible failure and the easiest to find.
  2. Put the answer near the top, in a passage that stands alone. Roughly 40 to 320 characters — long enough to be complete, short enough to quote whole.
  3. Be specific enough to be worth quoting. Replace one adjective with one number and the page becomes more citable than any amount of markup would make it.
  4. Find the third-party pages being cited for your category, and get named on them. Slower, unglamorous, and the highest-leverage item on this list.
  5. Measure as a rate, never as an event. Because the sources rotate, being cited once is not a result. Cited in nine of twelve weeks is.

The structural half of this is checkable on any URL for free with the page audit, and the fuller treatment is in the AEO guide.

Questions people ask

Does ranking first get you cited?

It helps and it does not decide it. Cited sources overlap heavily with pages that rank well, but the overlap is partial — pages from further down get cited, and pages ranking first frequently are not. Ranking is a strong prior, not a mechanism.

Why do forums and Reddit get cited so often?

Because a model writing a balanced answer needs several viewpoints, and a discussion thread supplies them in one document. A vendor page supplies one viewpoint and reads like advertising. This is uncomfortable for brands because it means a large share of your visibility lives on pages you cannot edit.

Can I pay to be cited?

No. There is no placement product in AI answers, no submission form and no index to join. Anyone offering this is selling something that does not exist — and the offer itself tells you what kind of vendor you are dealing with.

Does structured data get me cited?

It helps a machine understand what your page is about and which entity it belongs to, which is worth doing for the same reasons it always was. It is not a switch. No evidence supports treating schema as a citation lever, and treating it as one leads to markup describing a page nobody wrote.

Why do the cited sources change between runs?

Because the answer is generated rather than stored. The model is given retrieved documents and writes from them, and both the retrieval and the writing vary. Two people asking the same question minutes apart can get different sources — which is why a single check tells you almost nothing.

What actually moves the needle?

In rough order: being reachable by the AI crawlers at all, answering the question outright near the top of the page, being specific enough to be worth quoting, and being named on the third-party pages engines already read. The last one is the highest leverage and the least like traditional SEO.

Find out where you actually stand

One real question, real AI engines, and the answer they gave — including who was named in it. No account, no card.