Getting cited by an AI assistant is not one ranking event. A page has to pass at least four gates: it must be crawlable, retrieved for the question, selected as a source, and useful enough to shape the answer.
That distinction matters because most advice jumps straight to formatting. It tells businesses to add FAQ schema, publish an llms.txt, or make every paragraph shorter. Those may help a machine understand or discover a page. None, by itself, proves that the page will be selected or used.
We re-analysed a public 2026 dataset covering 23,745 citation records from ChatGPT, Google AI Overview/Gemini and Perplexity. Of those, 18,151 source pages were fetched successfully and could be compared. The pattern is useful, but narrower than a ranking-factor study: pages that influenced answers more were more relevant, more structured, and more likely to contain definitions, numbers, comparisons and procedures. That is association, not causation.
Here is what a business can responsibly do with that finding.
The four gates between publishing and citation
Treat AI visibility as a pipeline, not a switch.
| Gate | What has to happen | What you can control |
|---|---|---|
| 1. Eligibility | The system can crawl, index or otherwise access the page | Robots rules, indexability, renderability, stable URLs |
| 2. Retrieval | The page is considered relevant to the question | Topic coverage, entity clarity, words buyers actually use |
| 3. Selection | The system chooses the page among candidate sources | Unique evidence, reputation, corroboration, freshness |
| 4. Absorption | The answer uses facts or structure from the page | Definitions, numbers, comparisons, steps, clear sections |
Google's own guidance for AI search starts with the same foundation as ordinary search: pages must be indexed and eligible to appear with a snippet. It recommends unique, non-commodity, first-hand content and clear headings rather than special AI-only markup. The Google Search Central guidance is unusually plain on this point.
Our earlier answer engine optimization benchmark measured the first gate across thirteen Nepal tour operators. This study is about the other end of the pipeline: once a page is already in the source set, what distinguishes a passing reference from material the answer actually uses?
What we measured
The public GEO Citation Lab dataset contains 602 prompts across commerce, technology, local, healthcare, finance and news tasks. It records citations returned by ChatGPT, Google AI Overview/Gemini and Perplexity, then extracts 72 attributes from the cited pages.
We downloaded the dataset on 13 August 2026, verified its SHA-256 hash, and kept only rows where the source page was fetched successfully and an influence score was present. That left:
| Platform | Usable cited pages |
|---|---|
| ChatGPT | 3,323 |
| 6,385 | |
| Perplexity | 8,443 |
| Total | 18,151 |
The study's influence_score combines how often and how early a source is referenced, how much of the answer it covers, and text overlap between source and answer. It does not measure whether the page ranked first, caused a sale, or would be selected again tomorrow.
We sorted each platform separately, then compared its bottom and top influence quartiles. Separating the platforms matters: ChatGPT cited fewer sources per answer but used each source more deeply on average, so one pooled comparison would mix platform behaviour with page characteristics.
The source dataset, hash, exact filters and reproduction script are recorded in our measurement file. The underlying research and code are public in the GEO Citation Lab repository, and the accompanying paper describes the full 602-prompt framework.
The strongest pattern was relevance, not formatting
Across all three platforms, semantic alignment moved with influence. Median question-to-source similarity was higher in the top quartile on every platform:
| Platform | Bottom quartile | Top quartile |
|---|---|---|
| ChatGPT | 0.348 | 0.480 |
| 0.155 | 0.528 | |
| Perplexity | 0.140 | 0.529 |
The independent LLM relevance rating showed the same direction. Its median moved from 3 to 4 for ChatGPT, 2 to 4 for Google, and 1 to 4 for Perplexity.
This is the least glamorous result and the most important: a page has to answer the actual question. A polished article about a neighbouring topic is not made citation-worthy by schema. A service page that never states who the service is for, what it includes, or how it differs gives a retrieval system little to match.
For a business, the practical move is to map each important buyer question to one page with a direct answer. Do not make one generic page carry ten unrelated intents.
High-influence sources were easier to take apart
The top quartile also had more visible structure within every platform.
| Platform | Median words, low to high | Median headings, low to high | Median paragraphs, low to high |
|---|---|---|---|
| ChatGPT | 364 → 1,092 | 0 → 5 | 10 → 26 |
| 21 → 1,668 | 0 → 11 | 2 → 36 | |
| Perplexity | 21 → 1,182 | 0 → 7 | 2 → 28 |
This does not mean “write 1,668 words.” Page length is entangled with depth, site type and the number of claims available to reuse. Padding a weak page only produces a longer weak page.
The defensible interpretation is that answers can draw from pages containing multiple labelled, self-contained information units. Headings expose those units. Paragraphs give each claim a boundary. Lists and tables make relationships explicit. The page is easier to retrieve in pieces without losing its meaning.
Use a heading when the reader's question changes. Put the answer in the first sentence below it. Keep the evidence and its date in the same section rather than hiding all sources at the bottom.
Definitions, numbers and comparisons appeared far more often
Content form produced the clearest visible gaps. These are the percentages of pages in each platform's bottom and top quartiles that contained the feature:
| Feature | ChatGPT low → high | Google low → high | Perplexity low → high |
|---|---|---|---|
| Definition | 32.0% → 61.1% | 5.5% → 75.3% | 3.5% → 61.8% |
| Comparison | 13.3% → 32.7% | 1.9% → 43.5% | 0.9% → 35.7% |
| How-to steps | 13.7% → 31.4% | 1.7% → 39.3% | 1.1% → 31.9% |
| Numerical facts | 66.4% → 73.0% | 24.7% → 78.6% | 22.7% → 69.3% |
These forms are useful because they carry answerable information:
- A definition fixes what an entity or concept means.
- A number supplies a claim that can be checked and attributed.
- A comparison supplies a decision boundary.
- A procedure supplies an ordered action.
By contrast, Q&A formatting itself did not separate the groups consistently. A page does not become useful because headings end with question marks. This is why we treat FAQPage schema as a representation of genuine questions already answered on the page, not a citation tactic.
Original evidence gives a business something worth citing
Structure makes information extractable. It does not make the information original.
Google asks for first-hand, unique content because assistants can already summarise the commodity version. A business earns a reason to be cited when it contributes something another source cannot truthfully claim:
- a dated benchmark with a defined sample;
- a price range with scope and exclusions;
- an implementation result, including what failed;
- a comparison based on a repeatable test;
- a named expert's first-hand procedure;
- a case study whose numbers can be traced to records.
Our llms.txt implementation guide did not merely restate the specification. It fetched twenty-five likely early adopters, recorded every response, and found that only two of eighteen valid implementations had shipped the discovery relation added in version 2. The measurement is what makes the page a potential source. The file itself only helps an agent navigate.
For most companies, the evidence is already inside operations: proposal turnaround, common support questions, delivery lead times, audit failure rates, inventory movement, onboarding duration. Publish the aggregate, define the population and date it. Remove anything confidential. The result is more credible than a borrowed industry statistic and harder to replace with a generic summary.
Authority helps selection, but it is not a substitute for an answer
Retrieval systems do not judge a page in isolation. The organisation behind it, links and mentions from other sites, consistent entity information, and agreement with reliable sources all affect whether a claim is safe to use.
That makes distribution part of the work. Send an original study to the organisations included in it. Give partners a stable URL they can reference. Keep the company name, people, location and service definitions consistent across the website and verified profiles. Correct inaccurate third-party listings.
But do not confuse domain reputation with usefulness. The dataset only contains pages that were already cited; it cannot tell us how many equally structured pages were never selected. Nor does it establish a domain-authority threshold. A recognised domain may enter the candidate set more often, while a precise page from a smaller business may contribute the better passage.
The practical target is both: a source-worthy page on an entity that can be corroborated.
What schema, robots rules and llms.txt actually do
Technical controls matter, but at different gates.
| Control | Primary job | What it does not prove |
|---|---|---|
robots.txt | Allows or denies crawler access | That a permitted page will be selected |
| Canonical and indexability | Establish the page eligible for search | That it is the best answer |
| Structured data | Clarifies entities and page meaning | A citation or ranking boost |
llms.txt | Offers a curated site map to agents that choose to read it | Adoption by assistants or ranking value |
| FAQPage markup | Represents visible questions and answers | That Q&A formatting is influential |
Start with the defect that can stop the pipeline. A blocked crawler, accidental noindex, broken canonical or client-only page can make every content improvement irrelevant. Then make the page useful. Add structured representations last and keep them identical to what a person can see.
A six-step citation audit for a business page
- Test eligibility. Fetch the rendered page, inspect robots rules, canonical, status code and indexability. Check the specific crawler policy rather than assuming “AI crawler” is one category.
- Name the question. Write the exact buyer question this URL answers. If the first screen does not answer it, fix the page before adding markup.
- Find the unique claim. Identify the first-hand fact, benchmark, method or decision rule that only this organisation can own.
- Make it extractable. Give definitions, figures, comparisons and steps their own labelled sections. Keep dates, units and sample sizes beside the claims.
- Make it verifiable. Link primary sources, state limitations, name the author and correct conflicting entity information elsewhere.
- Measure repeatedly. Run a fixed prompt set across platforms and dates. Record whether the brand was mentioned, linked, cited and materially used. One screenshot is an observation, not a trend.
This is the same evidence discipline we use in SEO and AI-powered search engagements: eligibility first, then retrieval, then citation and business outcomes. The website development work supports those stages; it cannot manufacture the underlying expertise.
What this study cannot tell us
The dataset is a static 2026 snapshot, not a controlled intervention. Its prompts are not a random sample of every business question. Fetch failures removed 5,594 records from page-level comparison. Some extracted features and semantic labels are model-generated. The influence score measures use within one answer and includes text-overlap components, so it is not an independent measure of commercial value.
Most importantly, every page in the feature table was already cited. We can compare shallow and deep use among selected sources; we cannot estimate the probability that any ordinary page will be selected in the first place.
That limit matches the broader research. A 2026 critical survey of generative engine optimization found that topical relevance is comparatively reproducible while many popular optimisation claims do not yet show stable, cross-platform causal effects. An ACL 2026 comparison of generative search systems also found substantial variation in source reliance and stability across systems and repeated runs.
So the conclusion is deliberately modest: publish pages that are eligible, directly relevant, independently useful and easy to verify. Those qualities are associated with deeper citation use, help human readers regardless, and do not depend on a trick surviving the next model release.
If you need that chain tested on your site, book a 30-minute call. We will show which gate fails, the evidence for it, and what can be fixed without pretending to guarantee a citation.



