GrowNexus — Your Growth, Our Strategy
ServicesIndustriesMarketsProcessWorkInsights
GrowNexus — Your Growth, Our Strategy

360° growth partner for Kathmandu, Birgunj and global markets.

  • Facebook
  • Instagram
  • LinkedIn
  • TikTok

Digital Growth

  • Branding & Creative
  • Design Services
  • SEO Optimization
  • Growth Strategy & Consulting
  • Business Automation & AI
  • Training & Certification
  • Social Media Marketing
  • Website & App Development

Activation & BTL

  • Activation & Experiential
  • Event Marketing
  • Retail & Shopper Marketing
  • Direct Marketing
  • Influencer & Community
  • Promotions & Campaigns
  • Trade Marketing (B2B)
  • Out-of-Home (OOH)
  • Print & Collateral
  • Corporate Gifting & Merch
  • PR & Brand Building

Industries

  • Hospitality
  • E-commerce
  • SaaS
  • Professional Services
  • Education
  • Real Estate
  • PR & Communications
  • News & Digital Publishers

Markets

  • Markets hub
  • Nepal
  • United States
  • United Kingdom
  • Australia
  • Europe
  • UAE
  • India

Resources

  • Insights
  • Process
  • Pricing approach
  • Contact

Growth insights, no fluff

Monthly notes on SEO, paid social, and AI-native marketing for teams that ship.

© 2026 GrowNexus. All rights reserved.

  • Privacy
  • Sitemap
  • Terms
  • Cookies
  • Accessibility Statement
Skip to content
GrowNexus — Your Growth, Our Strategy
WorkInsightsAbout
Book a strategy call
AI & Automation
August 13, 2026·12 min read

What Gets a Business Cited by AI Assistants? 18,151 Sources Compared

We re-analysed 18,151 cited pages across ChatGPT, Google and Perplexity. Relevance, structure and evidence were associated with deeper use, but they are not magic ranking factors.

A field of source-page cards passing through four narrowing citation gates, with a small set of evidence-rich pages reaching an AI answer panel

Article body

On this page
  • The four gates between publishing and citation
  • What we measured
  • The strongest pattern was relevance, not formatting
  • High-influence sources were easier to take apart
  • Definitions, numbers and comparisons appeared far more often
  • Original evidence gives a business something worth citing
  • Authority helps selection, but it is not a substitute for an answer
  • What schema, robots rules and llms.txt actually do
  • A six-step citation audit for a business page
  • What this study cannot tell us

On this page

  • The four gates between publishing and citation
  • What we measured
  • The strongest pattern was relevance, not formatting
  • High-influence sources were easier to take apart
  • Definitions, numbers and comparisons appeared far more often
  • Original evidence gives a business something worth citing
  • Authority helps selection, but it is not a substitute for an answer
  • What schema, robots rules and llms.txt actually do
  • A six-step citation audit for a business page
  • What this study cannot tell us

The AI search series

One measurement programme, reported in three parts: what the field actually does, what the spec actually says, and what the citation data actually shows.

  1. Answer Engine Optimization: What 13 Measured Websites Actually Do
  2. How to Write an llms.txt: A Guide to v2, With 25 Sites Measured
  3. 3What Gets a Business Cited by AI Assistants? 18,151 Sources ComparedYou are here

Getting cited by an AI assistant is not one ranking event. A page has to pass at least four gates: it must be crawlable, retrieved for the question, selected as a source, and useful enough to shape the answer.

That distinction matters because most advice jumps straight to formatting. It tells businesses to add FAQ schema, publish an llms.txt, or make every paragraph shorter. Those may help a machine understand or discover a page. None, by itself, proves that the page will be selected or used.

We re-analysed a public 2026 dataset covering 23,745 citation records from ChatGPT, Google AI Overview/Gemini and Perplexity. Of those, 18,151 source pages were fetched successfully and could be compared. The pattern is useful, but narrower than a ranking-factor study: pages that influenced answers more were more relevant, more structured, and more likely to contain definitions, numbers, comparisons and procedures. That is association, not causation.

Here is what a business can responsibly do with that finding.

The four gates between publishing and citation

Treat AI visibility as a pipeline, not a switch.

GateWhat has to happenWhat you can control
1. EligibilityThe system can crawl, index or otherwise access the pageRobots rules, indexability, renderability, stable URLs
2. RetrievalThe page is considered relevant to the questionTopic coverage, entity clarity, words buyers actually use
3. SelectionThe system chooses the page among candidate sourcesUnique evidence, reputation, corroboration, freshness
4. AbsorptionThe answer uses facts or structure from the pageDefinitions, numbers, comparisons, steps, clear sections

Google's own guidance for AI search starts with the same foundation as ordinary search: pages must be indexed and eligible to appear with a snippet. It recommends unique, non-commodity, first-hand content and clear headings rather than special AI-only markup. The Google Search Central guidance is unusually plain on this point.

Our earlier answer engine optimization benchmark measured the first gate across thirteen Nepal tour operators. This study is about the other end of the pipeline: once a page is already in the source set, what distinguishes a passing reference from material the answer actually uses?

What we measured

The public GEO Citation Lab dataset contains 602 prompts across commerce, technology, local, healthcare, finance and news tasks. It records citations returned by ChatGPT, Google AI Overview/Gemini and Perplexity, then extracts 72 attributes from the cited pages.

We downloaded the dataset on 13 August 2026, verified its SHA-256 hash, and kept only rows where the source page was fetched successfully and an influence score was present. That left:

PlatformUsable cited pages
ChatGPT3,323
Google6,385
Perplexity8,443
Total18,151

Bar chart of the table above: usable cited pages analysed per platform, 3,323 for ChatGPT, 6,385 for Google and 8,443 for Perplexity.

The study's influence_score combines how often and how early a source is referenced, how much of the answer it covers, and text overlap between source and answer. It does not measure whether the page ranked first, caused a sale, or would be selected again tomorrow.

We sorted each platform separately, then compared its bottom and top influence quartiles. Separating the platforms matters: ChatGPT cited fewer sources per answer but used each source more deeply on average, so one pooled comparison would mix platform behaviour with page characteristics.

The source dataset, hash, exact filters and reproduction script are recorded in our measurement file. The underlying research and code are public in the GEO Citation Lab repository, and the accompanying paper describes the full 602-prompt framework.

The strongest pattern was relevance, not formatting

Across all three platforms, semantic alignment moved with influence. Median question-to-source similarity was higher in the top quartile on every platform:

PlatformBottom quartileTop quartile
ChatGPT0.3480.480
Google0.1550.528
Perplexity0.1400.529

Dumbbell chart of the table above: median question-to-source similarity for bottom and top influence quartiles on each platform. All three top-quartile values land close together.

The independent LLM relevance rating showed the same direction. Its median moved from 3 to 4 for ChatGPT, 2 to 4 for Google, and 1 to 4 for Perplexity.

This is the least glamorous result and the most important: a page has to answer the actual question. A polished article about a neighbouring topic is not made citation-worthy by schema. A service page that never states who the service is for, what it includes, or how it differs gives a retrieval system little to match.

For a business, the practical move is to map each important buyer question to one page with a direct answer. Do not make one generic page carry ten unrelated intents.

High-influence sources were easier to take apart

The top quartile also had more visible structure within every platform.

PlatformMedian words, low to highMedian headings, low to highMedian paragraphs, low to high
ChatGPT364 → 1,0920 → 510 → 26
Google21 → 1,6680 → 112 → 36
Perplexity21 → 1,1820 → 72 → 28

This does not mean “write 1,668 words.” Page length is entangled with depth, site type and the number of claims available to reuse. Padding a weak page only produces a longer weak page.

The defensible interpretation is that answers can draw from pages containing multiple labelled, self-contained information units. Headings expose those units. Paragraphs give each claim a boundary. Lists and tables make relationships explicit. The page is easier to retrieve in pieces without losing its meaning.

Use a heading when the reader's question changes. Put the answer in the first sentence below it. Keep the evidence and its date in the same section rather than hiding all sources at the bottom.

Definitions, numbers and comparisons appeared far more often

Content form produced the clearest visible gaps. These are the percentages of pages in each platform's bottom and top quartiles that contained the feature:

FeatureChatGPT low → highGoogle low → highPerplexity low → high
Definition32.0% → 61.1%5.5% → 75.3%3.5% → 61.8%
Comparison13.3% → 32.7%1.9% → 43.5%0.9% → 35.7%
How-to steps13.7% → 31.4%1.7% → 39.3%1.1% → 31.9%
Numerical facts66.4% → 73.0%24.7% → 78.6%22.7% → 69.3%

These forms are useful because they carry answerable information:

  • A definition fixes what an entity or concept means.
  • A number supplies a claim that can be checked and attributed.
  • A comparison supplies a decision boundary.
  • A procedure supplies an ordered action.

By contrast, Q&A formatting itself did not separate the groups consistently. A page does not become useful because headings end with question marks. This is why we treat FAQPage schema as a representation of genuine questions already answered on the page, not a citation tactic.

Original evidence gives a business something worth citing

Structure makes information extractable. It does not make the information original.

Google asks for first-hand, unique content because assistants can already summarise the commodity version. A business earns a reason to be cited when it contributes something another source cannot truthfully claim:

  • a dated benchmark with a defined sample;
  • a price range with scope and exclusions;
  • an implementation result, including what failed;
  • a comparison based on a repeatable test;
  • a named expert's first-hand procedure;
  • a case study whose numbers can be traced to records.

Our llms.txt implementation guide did not merely restate the specification. It fetched twenty-five likely early adopters, recorded every response, and found that only two of eighteen valid implementations had shipped the discovery relation added in version 2. The measurement is what makes the page a potential source. The file itself only helps an agent navigate.

For most companies, the evidence is already inside operations: proposal turnaround, common support questions, delivery lead times, audit failure rates, inventory movement, onboarding duration. Publish the aggregate, define the population and date it. Remove anything confidential. The result is more credible than a borrowed industry statistic and harder to replace with a generic summary.

Authority helps selection, but it is not a substitute for an answer

Retrieval systems do not judge a page in isolation. The organisation behind it, links and mentions from other sites, consistent entity information, and agreement with reliable sources all affect whether a claim is safe to use.

That makes distribution part of the work. Send an original study to the organisations included in it. Give partners a stable URL they can reference. Keep the company name, people, location and service definitions consistent across the website and verified profiles. Correct inaccurate third-party listings.

But do not confuse domain reputation with usefulness. The dataset only contains pages that were already cited; it cannot tell us how many equally structured pages were never selected. Nor does it establish a domain-authority threshold. A recognised domain may enter the candidate set more often, while a precise page from a smaller business may contribute the better passage.

The practical target is both: a source-worthy page on an entity that can be corroborated.

What schema, robots rules and llms.txt actually do

Technical controls matter, but at different gates.

ControlPrimary jobWhat it does not prove
robots.txtAllows or denies crawler accessThat a permitted page will be selected
Canonical and indexabilityEstablish the page eligible for searchThat it is the best answer
Structured dataClarifies entities and page meaningA citation or ranking boost
llms.txtOffers a curated site map to agents that choose to read itAdoption by assistants or ranking value
FAQPage markupRepresents visible questions and answersThat Q&A formatting is influential

Start with the defect that can stop the pipeline. A blocked crawler, accidental noindex, broken canonical or client-only page can make every content improvement irrelevant. Then make the page useful. Add structured representations last and keep them identical to what a person can see.

A six-step citation audit for a business page

  1. Test eligibility. Fetch the rendered page, inspect robots rules, canonical, status code and indexability. Check the specific crawler policy rather than assuming “AI crawler” is one category.
  2. Name the question. Write the exact buyer question this URL answers. If the first screen does not answer it, fix the page before adding markup.
  3. Find the unique claim. Identify the first-hand fact, benchmark, method or decision rule that only this organisation can own.
  4. Make it extractable. Give definitions, figures, comparisons and steps their own labelled sections. Keep dates, units and sample sizes beside the claims.
  5. Make it verifiable. Link primary sources, state limitations, name the author and correct conflicting entity information elsewhere.
  6. Measure repeatedly. Run a fixed prompt set across platforms and dates. Record whether the brand was mentioned, linked, cited and materially used. One screenshot is an observation, not a trend.

This is the same evidence discipline we use in SEO and AI-powered search engagements: eligibility first, then retrieval, then citation and business outcomes. The website development work supports those stages; it cannot manufacture the underlying expertise.

What this study cannot tell us

The dataset is a static 2026 snapshot, not a controlled intervention. Its prompts are not a random sample of every business question. Fetch failures removed 5,594 records from page-level comparison. Some extracted features and semantic labels are model-generated. The influence score measures use within one answer and includes text-overlap components, so it is not an independent measure of commercial value.

Most importantly, every page in the feature table was already cited. We can compare shallow and deep use among selected sources; we cannot estimate the probability that any ordinary page will be selected in the first place.

That limit matches the broader research. A 2026 critical survey of generative engine optimization found that topical relevance is comparatively reproducible while many popular optimisation claims do not yet show stable, cross-platform causal effects. An ACL 2026 comparison of generative search systems also found substantial variation in source reliance and stability across systems and repeated runs.

So the conclusion is deliberately modest: publish pages that are eligible, directly relevant, independently useful and easy to verify. Those qualities are associated with deeper citation use, help human readers regardless, and do not depend on a trick surviving the next model release.

If you need that chain tested on your site, book a 30-minute call. We will show which gate fails, the evidence for it, and what can be fixed without pretending to guarantee a citation.

Cite this article

Arman Ahamed (2026). "What Gets a Business Cited by AI Assistants? 18,151 Sources Compared." GrowNexus, August 13, 2026. https://grow-nexus.com/insights/how-to-get-cited-by-ai-assistants/

Share on LinkedIn

Written by

Arman Ahamed

Digital Marketing Manager, GrowNexus

Leads growth strategy and AI-native marketing programs for Nepal and global clients. Writes about what ships, not what slides.

Who this is for

  • E-commerce marketing
  • SaaS marketing
  • PR & Communications marketing
  • News & Digital Publishers marketing
  • Plumbing & HVAC marketing

Keep reading

Related insights

More from the same topic, selected by category overlap.

AI & Automation
Aug 22, 2026/8 min read

How to Format a Press Release So AI Assistants Cite It

One checkable fact, a named human, real questions, and the right schema. What actually makes a release extractable in 2026 — including the FAQ markup change most advice has not caught up with.

AI & Automation
Aug 22, 2026/10 min read

How to Measure AI Citations from PR When Your Analytics Show Nothing

Assistant referrals look like a rounding error while page-one buying questions earn no clicks. Four measurement methods, what each one misses, and the three report splits that change what you work on next.

AI & Automation
Aug 13, 2026/8 min read

How to Write an llms.txt: A Guide to v2, With 25 Sites Measured

The llms.txt spec was revised in August 2026. We measured 25 major sites on 13 August: 18 publish a file, but only 2 have implemented the discoverability relations v2 added, and none ship them in HTML. Here is how to write one properly.

Strategy notes, no fluff

One email when we publish something worth your time. Unsubscribe anytime.