Skip to content

SEO and AI search

How to get your business cited by ChatGPT

ChatGPT does not rank the web the way Google does. It retrieves a shortlist, mostly from Bing, then decides which of those pages are clear and credible enough to quote. That is two separate gates, and most businesses fail the first one without knowing it, because they have never checked whether Bing has them at all.

K.M. Abdullah Probal

Websites, tracking, automation, and technical search

11 min read · Published September 7, 2026

Research notes with one passage marked in red beside a stack of discarded reference pages

How does ChatGPT actually pick sources?

In two stages. It retrieves candidate pages from a live search index, primarily Bing, then a separate selection step decides which of those retrieved pages are worth quoting. Being retrieved and being cited are different outcomes.

Almost every mistake in this area comes from treating it as one gate. People optimize a page the way they would for Google, watch nothing happen, and conclude the whole thing is unknowable. It is not unknowable. It is just two problems wearing one name.

Stage one is retrieval and it is largely an index question. If Bing does not have your page, no amount of clever writing puts you in the answer. Stage two is selection, and that is a quality and clarity question about the pages that did get retrieved.

Third-party studies of ChatGPT citations suggest only a small fraction of retrieved pages actually end up quoted, with the rest evaluated and discarded silently. Treat the exact percentages loosely, since methodologies vary and none of them are OpenAI's own numbers. The shape is the useful part. Retrieval is necessary and nowhere near sufficient.

Why does Bing matter so much here?

Because it supplies most of the retrieval layer. B2B teams have spent a decade ignoring Bing as a traffic source, which was reasonable, and that habit now quietly blocks them from an AI surface that matters.

Check this before you change anything else. Search your own brand and a couple of your key pages on Bing directly. If they are missing or thin, that is your bottleneck, and it is a much cheaper fix than rewriting your content library.

Bing Webmaster Tools will show you coverage and let you submit URLs. There is also IndexNow, a keyless push protocol where you host a key file at your site root and POST a list of URLs. Bing, Yandex, Seznam, and Naver all read from the same submission. We use it on this site, and it costs nothing to run.

None of this touches Google. Google does not participate in IndexNow and its index is separate. So this is an addition to your search work, not a replacement for it.

What makes a page quotable rather than merely relevant?

A passage that answers the question directly, early, and can survive being lifted out of the page without losing meaning. Selection favors clarity over comprehensiveness.

Think about what the model is doing. It needs a span of text it can attribute to you that stands on its own. A paragraph that only makes sense after three paragraphs of setup is unusable, no matter how good it is.

Which is why the long throat-clearing introduction is so expensive here. Every sentence before your actual answer is a sentence that pushes the quotable part further from the top.

  • Answer the question in the first third of the page, not the last
  • Write self-contained passages that carry their own context
  • Use the question as a heading, so the match is unambiguous
  • Include something original: a number you measured, a case you handled, a limitation you found
  • Keep facts consistent across your site, because contradictions make a source risky to quote
  • Date the page and keep it current, since recency weighs in selection

Does structured data help?

Indirectly. Schema clarifies what a page is and keeps your entity facts consistent, which supports selection. But no schema type forces a citation, and Google has said no special markup is required for its AI features.

Be skeptical of anyone selling a schema type as an AI ranking lever. Structured data that matches visible content is good practice and worth doing. It is not a switch.

The higher-value structural work is plainer than that. Clear headings phrased as questions. Direct answers under them. Consistent naming for your company, your services, and your people everywhere those facts appear.

What if we are small and new?

You are working against a real bias. Selection tends to favor sources already well represented in training data, which means Wikipedia, Reddit, and large publishers. A new domain cannot out-authority them, so it has to out-specify them.

The opening is narrowness. Big publishers write the general answer because that is where their volume is. They will not write the awkward, specific, operational question your buyers actually ask, because on its own it looks too small to bother with.

That question is where you win. Not because you outrank anyone, but because when someone asks it, there may be very little else genuinely useful for the model to retrieve.

The other lever is corroboration off your own site. Being described consistently in places that already carry weight, industry directories, review platforms, communities where your buyers actually talk, does more for a young brand than another page on its own domain.

Do we need to allow AI crawlers?

Yes, if you want to be quotable. Check robots.txt for OAI-SearchBot, GPTBot, PerplexityBot, ClaudeBot, and Google-Extended. Blocking them is a legitimate business decision, but it forfeits this channel.

Worth understanding that these bots do different jobs. Some fetch pages to answer a live question, others gather training data. A blanket block treats them as one thing and usually reflects a decision nobody consciously made, often inherited from a template or a security plugin.

Read your own robots.txt today. We find blocks in it that surprise the people who own the site more often than you would expect.

How do you measure any of this?

Fix a prompt set and re-run it on a schedule. Record whether you appear, which URL was cited, and whether what was said about you is accurate. Accuracy matters as much as presence.

Pick twenty to forty questions a real buyer would type. Run them, log the result, repeat monthly. Answers vary between runs and change over time, so a single check tells you almost nothing while a tracked set tells you a direction.

Watch for the failure that gets missed: being mentioned inaccurately. Wrong services, wrong location, a competitor's claim attached to your name. That is worse than absence and it usually traces back to inconsistent facts across your own pages and profiles.

Then connect it to the business. Referral sessions from AI surfaces, what those visitors do next, and whether any of it becomes pipeline. Mention counts are not revenue and should never be reported as if they were.

What can nobody promise you?

A citation. Selection is controlled by the platform, varies by query, changes between runs, and has no submission process. Anyone guaranteeing placement in an AI answer is describing something they do not control.

What is genuinely controllable: whether you are in the index, whether crawlers can reach you, whether your best answer is near the top of the page, whether your facts agree with each other, and whether anything on your site is worth quoting at all.

That list is unglamorous and it is most of the work.

Questions buyers ask

Direct answers for the questions that usually appear before a buying decision.

Is getting cited by ChatGPT the same as ranking on Google?+

No. ChatGPT retrieves mostly from Bing and then runs a separate selection step. A page can rank well on Google and never be cited, and the reverse happens too.

Do we need an llms.txt file?+

It can communicate preferences to tools that choose to read it, and it costs little. It is not a ranking instruction and no major system treats it as one. Publish it if you like, but do not expect it to move anything on its own.

How long before we see citations?+

Indexing changes can show within weeks. Selection depends on whether anything you publish is genuinely the best available answer, which is a content problem rather than a waiting problem.

Can we pay to appear in ChatGPT answers?+

There is no citation placement product to buy. Be suspicious of anyone offering one.

ChatGPT says something wrong about our company. What now?+

Fix the underlying facts first, across your own site, your schema, and your third-party profiles. Inconsistent public facts are the usual source, and correcting them is the only durable lever you hold.

Need help applying this to your business? See Answer Engine Optimization.