How do AI answers find and cite sources?
How query fan-out, retrieval, grounding, source selection, and citations work, and why an AI answer has no single stable ranking.
Sascha KirsteinIn this article
- The short version
- Model knowledge or live retrieval
- Step 1: prompt interpretation and query fan-out
- Step 2: retrieval and grounding
- Step 3: source and passage selection
- Step 4: synthesis
- Step 5: citations
- Why there is no single LLM ranking
- How to measure source visibility responsibly
- What publishers can control
AI answers do not usually come from one ranked list. A system may answer from information learned during training, retrieve current sources, or combine both. When it uses search, the path often includes query reformulation, retrieval, passage selection, answer generation, and visible citations.
That sequence explains why a page can rank well in search but remain absent from an AI answer. It also explains why a cited page is not automatically the answer's most important source.
The short version
A search-connected AI system commonly does five things:
- interprets the prompt and may turn it into several searches;
- retrieves pages, documents, or structured data;
- selects passages that appear useful for the question;
- combines claims from the selected material into an answer;
- surfaces some supporting sources as citations.
Providers do not publish one common recipe, and the details change. Google, ChatGPT, Copilot, and Perplexity can return different sources for the same question. Country, language, date, account context, and previous messages can change the result too.
Model knowledge or live retrieval
The first distinction is whether the answer uses current sources at all.
A language model can generate an answer from patterns learned during training. In that case, the user may see no source list and cannot trace a sentence to one current webpage. The information may still reflect material that appeared in the training data, but that is different from retrieving and citing a page for this answer.
Search-connected products can fetch current information. OpenAI describes ChatGPT search as a product that uses search providers and partner content. Google says AI Overviews and AI Mode use its core Search systems. A single response may mix retrieved material with the model's existing knowledge.
This distinction matters for measurement. A visible citation proves that the product surfaced a source with that answer. No citation does not prove that a topic or brand had no influence on the model.
Step 1: prompt interpretation and query fan-out
The user's prompt is not always the search query.
For a broad question, the system may identify subtopics and issue several related searches. Google calls this query fan-out. Its documentation says AI features can run multiple searches across subtopics and data sources to build a response.
Suppose someone asks: "Which PR analytics tools are suitable for an international communications team?"
The system could search separately for product capabilities, supported countries, reporting features, pricing, reviews, or recent comparisons. These internal searches may retrieve pages that do not rank for the user's exact wording.
Query fan-out is a documented Google term, not a universal name for every provider's process. Microsoft described the new Bing's orchestrator as creating and refining internal queries. OpenAI says ChatGPT search can rewrite prompts into targeted searches, but it does not publish a universal query-generation formula.
Bing gives verified site owners a limited view of this process. Its AI Performance report includes sampled grounding queries connected to cited pages. These phrases are useful first-party evidence of retrieval behavior. They are a sample, not a complete query log or ranking report.
Step 2: retrieval and grounding
Retrieval collects current pages, passages, or data that might help answer the question. Grounding connects the generated answer to that material.
Search quality still matters at this stage. A page must be accessible to the relevant crawler or search system, discoverable, and relevant enough to enter the candidate set. Crawler access creates eligibility. It does not guarantee retrieval or citation.
Google states that the same technical requirements and search policies apply to its generative search features. There is no special AI schema that guarantees inclusion. OpenAI likewise separates OAI-SearchBot controls for ChatGPT search from the controls used for model training.
Grounding reduces the need to answer only from stored model knowledge. It does not make the result infallible. The system can select an old page, misunderstand a passage, combine incompatible claims, or attach an incomplete citation.
Step 3: source and passage selection
Retrieval can return many candidates. The system then decides which material to use.
The useful unit is often a passage, not a whole page. One paragraph may supply a product fact, another source may supply a price, and a third may support a market comparison. A page can contribute one sentence without becoming the main source for the response.
Selection may depend on relevance, freshness, language, location, source quality, and how directly a passage supports the claim. The providers do not expose a complete, stable weighting formula. Claims about a fixed list of "LLM ranking factors" therefore go beyond the available evidence.
Clear source material still helps. Specific facts, named methods, current dates, primary evidence, and consistent entity information are easier to retrieve and interpret than vague claims. Independent reporting, reviews, public records, and expert sources can matter alongside a company's own pages.
Step 4: synthesis
The model turns selected material into a new answer. It may summarize, compare, quote briefly, or combine several sources in one paragraph.
This is where AI discovery differs most visibly from a classic search result. The user receives a composed response instead of ten blue links. Brands may be:
- named or omitted;
- recommended, compared, or criticized;
- described accurately or incorrectly;
- supported by owned or third-party sources;
- visible without receiving a click.
The answer can also change across repeated runs even when the prompt stays the same. Generated wording is variable, and retrieval results can change as indexes and source pages change.
Step 5: citations
A citation is a source link that the provider surfaces with an answer. It gives the user a way to inspect supporting material.
A citation does not prove that:
- the page has a universal AI rank;
- the page was the strongest influence on the answer;
- every claim in the sentence came from that page;
- the answer is accurate;
- the source will appear again for the same prompt.
Citation order is also not equivalent to search position. One source may support a minor factual detail while another shapes most of the comparison. Some products group several citations after a paragraph, which makes claim-level attribution harder.
Microsoft makes this distinction explicit in its AI Performance reporting. Citation counts and cited pages describe visibility in grounded answers. They do not reveal a page's authority score or precise role in generating a response.
Why there is no single LLM ranking
"Where do we rank in ChatGPT?" sounds like an SEO question, but it hides several different measurements.
A useful report separates them:
| Question | Suitable observation |
|---|---|
| Does the brand appear? | Mention or visibility rate |
| How early and substantially does it appear? | Position and prominence |
| Which other brands appear? | Share of Voice |
| Which sources are surfaced? | Citation rate and source mix |
| Is the description correct? | Message accuracy and factual review |
| Is the framing favorable? | Brand-specific sentiment or framing |
None of these is a permanent rank. Each describes a sample of answers collected under defined conditions.
How to measure source visibility responsibly
Start with a governed set of prompts that reflects real stakeholder questions. Record the provider, model or product, prompt, country, language, date, and relevant settings. Repeat important prompts because one run can be an outlier.
For each answer, save:
- the brands mentioned and their position;
- the cited URLs and domains;
- the claim each citation appears to support;
- whether key facts and messages are correct;
- the answer's framing;
- referrals or conversions when a click occurs.
As documented in August 2026, Google, Bing, and OpenAI expose different first-party observations. Google reports eligible generative search impressions in Search Console. Bing reports citations and cited pages in Webmaster Tools. OpenAI adds utm_source=chatgpt.com to ChatGPT search referrals. These signals cannot be merged into one exact cross-platform rank.
The AMEC GEO Principles recommend transparent prompt samples, repeated measurement, and clear limits. They also warn against presenting one tool or score as complete AI visibility.
What publishers can control
Publishers cannot control the final answer. They can improve the information available to the system:
- keep important pages crawlable and indexable;
- publish specific, current, verifiable facts;
- state methods, definitions, and limitations plainly;
- keep names and claims consistent across reliable sources;
- earn independent coverage and expert references where they add real evidence;
- measure recurring prompts instead of chasing one screenshot.
This work is part of Generative Engine Optimization. The distinction from conventional search is covered in GEO vs. SEO. For the measurement system, continue with What is AI visibility? or read aclipp's exact AI Visibility metrics and source classifications.
Keep reading
Insights
How B2B Companies Communicate on Social Media
Almost 800 communicators were surveyed about their approach. The results are fascinating and summarized here.

Insights
How is AI Changing the Day-to-Day of PR?
AI-based software is already significantly transforming the PR industry today. But where are the greatest potentials for PR hidden?

Knowledge
How Can You Measure the Value of Social Media Coverage?
How can you evaluate the performance of social media clippings across platforms and channels? We provide expert tips!
