Skip to content

AI Search in 2026: How to Get Cited by ChatGPT, Gemini, Google AI Mode, Perplexity, and Copilot

AI Search in 2026: How to Get Cited by ChatGPT, Gemini, Google AI Mode, Perplexity, and Copilot

To improve your chance of being cited in AI search, make each important page accessible to the engines you want to reach, answer a specific question clearly, support claims with original evidence or primary sources, and keep your company and product facts consistent across the web. No step guarantees a citation. Retrieval is query-dependent, and the major vendors do not publish a universal source-selection formula.

For Google Search, the technical answer is straightforward: a page must be indexed and eligible to appear with a snippet. Google’s guide to AI features in Search says existing SEO practices still apply and that llms.txt, special AI markup, and AI-specific schema are not required. Other engines have their own crawler controls, which are covered below.

Use four controllable inputs as your working model:

Input What to verify What it can improve
Access The desired crawler is allowed by robots.txt, the CDN, and the WAF, and receives the page content Whether the engine can retrieve or index the page
Answer quality The page gives a direct, specific answer and makes the supporting evidence easy to understand Whether the page is useful for the current query
Evidence Claims have first-party data, firsthand experience, or current primary sources Whether the answer is supportable and distinct from a summary
Entity consistency Names, product facts, author details, and business information agree across owned and credible third-party sources Whether the source and its claims are easy to identify and corroborate

Improving all four raises eligibility. It still does not ensure that an engine will retrieve, quote, or link the page for a particular prompt.

What AI search is, and how it changes click behavior

Classic SEO focuses on visibility in ranked search results. Answer engines may synthesize a response, quote or paraphrase sources, and offer only a small set of links. AEO (answer engine optimization) and GEO (generative engine optimization) describe the work of improving visibility inside those answers. For Google Search specifically, Google says the relevant practices remain part of SEO.

Ahrefs re-ran its CTR study on December 2025 data and published the update on February 4, 2026. In that sample, the presence of an AI Overview correlated with a 58% lower average click-through rate for the top-ranking page. That is a correlation from one dataset, not a forecast for every site or query. For a longer view of the shift, see our retrospective on what AI Mode’s first year did to organic traffic.

At Google I/O 2026, Sundar Pichai said AI Overviews had passed 2.5 billion monthly users and AI Mode had passed 1 billion. At that scale, marketers should check whether the pages tied to buyers’ questions can be discovered, understood, and supported well enough to be considered as sources.

How each major engine chooses and cites sources

The engines are not interchangeable. They differ in what they crawl, how they retrieve, and where citations even appear. Treat them separately.

Diagram showing source pages, profiles, videos, and news references feeding into an AI answer with citations
Each answer engine retrieves from a different mix of indexes, sources, and supporting context before it decides what to cite.

Google AI Overviews and AI Mode

Google says AI Overviews and AI Mode use the Search index and may use query fan-out to run several related searches. A page must be indexed and eligible to appear with a snippet. Existing fundamentals still apply: crawlability, internal links, useful text and media, and current Business Profile or Merchant Center data where relevant. Google also notes that meeting the requirements does not guarantee inclusion. On June 3, 2026, Google began rolling out dedicated generative-AI performance reports to a subset of Search Console properties. If the report is available in yours, it covers AI Overviews, AI Mode, and generative AI features in Discover.

Ahrefs analyzed 863,000 SERPs and 4 million AI Overview URLs in March 2026 and found that 37.9% of cited URLs ranked in the top 10 for the exact query. Another 31.2% ranked from positions 11 to 100, and 31% were outside the top 100. The study shows that citation sources do not have to rank in the top 10 for the exact query; it does not show that answering an adjacent question will cause a citation. YouTube represented 5.6% of all citations in the sample. In late May 2026, Google also brought Preferred Sources into AI Overviews and AI Mode, allowing users to give selected publishers more prominence in eligible results.

OpenAI says ChatGPT Search gives timely answers with links, searches automatically when it judges the web would help, and also lets users trigger search manually. Responses can carry inline citations; when they do not, users can open a Sources panel.

For publishers, access is the first technical check. OpenAI’s publisher FAQ says pages should allow OAI-SearchBot if they want content included in ChatGPT search summaries and snippets. OpenAI also says hosts and CDNs should allow its published crawler IP ranges, not only the user agent. GPTBot is a separate control for potential model training; blocking it does not by itself remove a site from ChatGPT search. Referral URLs include utm_source=chatgpt.com, which makes those visits identifiable in analytics.

OpenAI has not published a detailed source-ranking model and does not promise top placement. Allowing the search crawler is an eligibility step, not a citation guarantee.

Perplexity

Perplexity’s robots.txt guidance says PerplexityBot respects robots.txt and is used for search-style indexing, not foundation-model training. If a site blocks the bot, Perplexity says it will not index the page’s full or partial text. It may still retain limited information such as the domain, headline, and a brief factual summary.

Perplexity does not publish a complete source-weighting formula. If you want the page text to be eligible, make sure PerplexityBot is not blocked accidentally. That does not ensure the page will be cited.

Google Gemini

Gemini Apps are less transparent than Google Search, but Google’s help docs clarify the user experience. Gemini may show sources and related links, including public websites, uploaded files, or connected Workspace content. Not every response includes sources. If Gemini quotes a large amount of text from a webpage, Google says it shows a link to that page in the sources list, and web images are linked to their sources.

On the developer side, Grounding with Google Search can return grounding metadata with web sources and citations. One crawler control is worth separating from Search. Google-Extended governs whether Google-crawled content may be used for future Gemini model training and for grounding in Gemini Apps and Vertex AI. Google states that this control does not affect Google Search inclusion or ranking.

Microsoft Copilot

For information-seeking chats, Microsoft says Copilot Chat and Agents can use web search by generating a short query from the prompt and sending it to Bing. The Sources button shows the query and sources used. Microsoft also documents that web-search citations are available in Microsoft 365 Copilot Chat, but not in the Copilot pane inside apps such as Word or PowerPoint. Bing crawl and index access is therefore a sensible eligibility check, but Microsoft does not publish a universal citation-ranking formula.

What the latest data actually says

These studies can help prioritize tests, but they are correlational and use different engines, prompts, and samples. None establishes a universal citation factor.

Semrush’s content optimization study, published January 14, 2026, reported positive correlations for clarity and summarization (+32.8%), E-E-A-T signals (+30.6%), Q&A format (+25.5%), section structure (+22.9%), and structured data (+21.6%). Its technical SEO study of 5 million cited URLs, published January 5, found Organization, Article, and BreadcrumbList among the most common schema types on cited pages. Common does not mean causal, and Google says special schema is not required for its AI features.

Muck Rack analyzed more than 25 million cited links from sampled ChatGPT, Claude, and Gemini responses across 17 industries in May 2026. Within that dataset, sources the researchers classified as earned media accounted for 84% of citations, and journalism accounted for 27%. That supports testing editorial coverage as part of an AI-visibility program; it does not mean earned media will account for 84% of citations in every industry or engine.

Previsible analyzed 1.96 million LLM-driven sessions from November 2024 through November 2025. AI referrals averaged 0.13% of total sessions across its sample and were concentrated on page types such as industry, tools, blog, and pricing pages. Use that as a benchmark from one dataset, not a universal traffic or intent profile. Measure referral volume, lead quality, and conversions on your own site.

How to improve citation eligibility across engines

Work through access, answer quality, evidence, and corroboration in that order. The checklist below turns those ideas into verifiable steps.

Layered citation eligibility stack showing crawler access, source evidence, structured data, and off-site references feeding a cited answer
Access, structure, evidence, schema, and off-site references can all support eligibility, but none guarantees selection.

A practical eligibility audit

Check How to verify it Low-risk action
Crawler access Review robots.txt, CDN and WAF rules, HTTP status codes, and server logs for the bots you want Remove accidental blocks; for OpenAI, also allow the published OAI-SearchBot IP ranges
Google eligibility Use URL Inspection and review noindex, canonical, and snippet controls Fix unintended indexing or snippet restrictions; do not remove deliberate privacy or legal controls
Visible answer Read the rendered page without navigation or supporting context Put a short, complete answer near the relevant heading and define any necessary terms
Evidence Check every dated, comparative, or technical claim Add original data, firsthand examples, or links to current primary documentation
Entity consistency Compare the site, author bio, profiles, schema, and authoritative listings Correct conflicting names, descriptions, dates, and product facts
Measurement Review Search Console reports, referral analytics, and a fixed set of representative prompts Save a dated baseline and track citations, qualified visits, and conversions over time

Fix crawler access and indexability first. Google AI features need the page indexed and snippet-eligible. ChatGPT search summaries need OAI-SearchBot access from OpenAI’s published IP ranges. PerplexityBot should not be blocked if you want Perplexity to index the page text. Check the page response and server logs rather than relying only on a user-agent test.

Publish information another page cannot reproduce from the same sources. Google’s guide recommends useful, non-commodity content. For a startup, strong candidates include original benchmarks, first-party data, implementation lessons, honest technical comparisons, real pricing context, and customer-pattern insights. Explain the method, sample, date, and limitations so a reader can evaluate the evidence.

Make the answer easy to understand and quote. Start the relevant section with a complete answer, then show the evidence and caveats. Use descriptive headings, comparison tables, and Q&A sections when they help the reader. Semrush found correlations between citations and clarity, Q&A formatting, and section structure, but Google explicitly says there is no required AI-specific format or text chunking. For a deeper content method, see our guide to turning blog posts into knowledge hubs.

Use structured data to describe visible facts. Google says special schema is not required for AI features. Use supported markup for the page type, keep it consistent with visible content, and validate it. Organization, Article, Breadcrumb, Product, and other appropriate types can clarify entities and page relationships; they are not citation switches. We review the evidence and limits in our guide to technical SEO for AI search.

Correct conflicting facts across the web. Keep company, product, founder, and category claims consistent across your website and schema, company and executive profiles, videos, earned media, customer stories, and Business Profile or Merchant Center where relevant. The Muck Rack sample shows that third-party editorial sources appear frequently in some AI citations. That makes accurate off-site coverage worth testing alongside owned content, without assuming that a mention will produce a link.

Refresh time-sensitive pages when the substance changes. Update examples, screenshots, vendor documentation, and dated evidence. Show an accurate modified date when the content has materially changed. Changing only the date does not make a page more useful or more likely to be cited.

What startup and B2B teams should do first

If you have limited time this quarter, sequence the work like this.

First, audit the pages closest to revenue. Confirm crawler access, add a direct answer where it helps the reader, check technical and dated claims against primary sources, and correct conflicting entity or author details. Use our AEO checklist for a page-by-page review.

Next, create assets that add evidence the existing results do not have: original data, technical comparisons with stated criteria, and firsthand implementation lessons. Publish in the formats your audience uses, which may include your site, video, or executive-led social content.

Last, track a fixed set of representative prompts, AI referral traffic, and downstream conversions. Record the engine, date, prompt, cited URL, and answer because results can change between runs. Optimize for qualified outcomes rather than citation count alone.

For teams that want help with the technical and editorial side together, Triaza’s AI Search & Answer Engine Optimization service pairs naturally with traditional Search Engine Optimization. If you need the baseline SEO foundation first, start with SEO Fundamentals for Startups or the deeper SEO and AEO guide.

Questions marketers keep asking

Do I need llms.txt?

Not for Google Search visibility. Google says you do not need machine-readable AI text files, Markdown files, or special markup to appear in its generative search features. The confusing part is Lighthouse: Chrome’s newer Agentic Browsing audit does check llms.txt, but Google also says that category is experimental, does not use the normal 0-100 Lighthouse score, and marks a missing file as not applicable rather than a failure.

You may publish one for tools that choose to use it, but do not treat it as a Google ranking factor or citation guarantee. A normal 404 for /llms.txt does not violate Google’s search guidance.

Links can help search engines discover pages and may support classic SEO, but no vendor documents a universal backlink formula for AI citations. In Muck Rack’s May 2026 sample, sources it classified as earned media made up 84% of citations. Treat that as a reason to test accurate editorial coverage, not as a target link ratio or a promise of citation visibility.

Why is my top-ranking page not cited?

One reason is that Google may retrieve sources through related searches rather than only from the visible top 10 for the exact wording. In Ahrefs’ March 2026 sample, 37.9% of cited URLs ranked in the top 10 for the exact query. That does not reveal why your page was omitted. Check eligibility, answer fit, evidence, freshness, and source consistency before choosing a fix.

Is AI search replacing SEO?

No. Google explicitly places its generative search features within existing SEO practices. Other engines add their own crawler and citation behavior, so AEO broadens the audit; it does not make crawlability, indexability, and useful content obsolete.

How do I measure AI visibility?

Combine three things: dated citation checks for a fixed prompt set, AI referral traffic in analytics, and conversions from those sessions. Do not assume AI traffic is always low-volume or high-intent. Previsible observed that pattern in its sample; your results may differ by market and page type.

Where to start this week

Start with three pages tied to qualified demand:

  1. Check robots.txt, CDN and WAF rules, HTTP responses, and server logs for Googlebot, OAI-SearchBot, and PerplexityBot. For OpenAI, verify access from its published IP ranges and decide separately whether to allow GPTBot for training.
  2. Use Google’s URL Inspection tool to confirm the page is indexable, canonicalized as intended, and not restricted from snippets. Use the generative-AI report if it is available in your Search Console property.
  3. Add a short, direct answer where it improves the page, then verify every technical or dated claim against a primary source. State the evidence’s date, sample, and limitations.
  4. Correct conflicting company, product, and author facts across the site, structured data, profiles, and authoritative listings.
  5. Save a representative prompt set and record citations, referral visits, and conversions monthly. Treat changes as observations, not proof that a single edit caused the result.

That process will not guarantee a citation. It will tell you whether the basic eligibility, evidence, and measurement work is in place before you invest in a larger content program.

Ready to put these insights into action?

Let's discuss how Triaza can help your business grow.

Talk with us

Resources

Go To The Blog