mechanics· · 10 min read

How AI answer engines decide which firm to recommend

A firm can be eliminated at any of seven steps before it ever reaches the answer. What decides the outcome at each step, which levers you actually control, and why most spending goes to the wrong one.

on this page

Executive summary

When someone asks ChatGPT for the best CA firm in Mumbai, the model does not consult a ranked list. It runs a sequence of seven steps, and a firm can be eliminated at any one of them without ever reaching the answer. Most firms fail at step four, retrieval, for a reason that has nothing to do with the quality of their website. This post walks through each step, names what decides the outcome at that step, and separates the levers you control from the ones you do not. The practical conclusion is that inclusion is a sequence of gates, so the order you fix things in matters more than how much you spend.

Why "the AI just knows things" is the wrong model

The most common assumption I hear from partners is that AI systems hold a directory of businesses and surface the good ones. They do not.

A model trained on text has a compressed, lossy memory of what appeared in that text. For Deloitte, that memory is dense. For a 12-partner practice in Fort, it is usually empty. So when the question is specific and local, the system goes and looks something up, in real time, using a search index. What it finds in that moment, and how that content is written, decides who gets named.

Being recommended by an AI is not a reputation outcome. It is a retrieval outcome. Those are different problems with different fixes.

The seven steps behind a single recommendation

01 — THE SELECTION PIPELINE Seven steps. Seven places to be eliminated. "best CA firm in Mumbai for GST" 01 Prompt interpretation Intent, location and constraints are extracted. Nothing you control. 02 Parametric recall What the model already knows. For most Indian firms, nothing. 03 Query fan-out One question becomes five to fifteen searches you never see. 04 Retrieval The hard gate. Not in the index means eliminated here, silently. 05 Reranking Candidates scored on relevance, authority and corroboration. 06 Grounding Specific passages selected. Format decides this, not quality. 07 Synthesis and naming The answer is written. Your firm is named, cited, or left out. Elimination at any step is silent. You never see the answer you did not appear in.

Step 1: prompt interpretation

The system reads intent, location and any constraints out of the question. "Best CA firm in Mumbai for GST" carries a service, a city and an implied quality judgment. Nothing here is under your control. It matters only because it sets up the next step.

Step 2: parametric recall

The model checks what it already holds. This is where large, heavily-written-about brands win without any retrieval at all.

For an Indian professional firm, assume this is empty. Parametric presence is built over years of being written about in public, at volume, on sources that end up in training data. It is a genuine long-term goal and a useless short-term one. Anyone selling you a three-month plan to get "into the model's knowledge" is selling you something that does not work that way.

Step 3: query fan-out

The system rewrites the question into multiple sub-queries and searches each one. "Best CA firm in Mumbai for GST" might become GST consultants Mumbai, chartered accountant GST filing BKC, top CA firms Mumbai reviews, GST compliance services near me, and several more.

This is the single most consequential mechanic for how the work gets planned. You are not competing for the query the client typed. You are competing for a family of rewritten queries you never see, and your coverage across that family determines whether you survive to the next step. A page optimised for one phrase covers one branch of a tree with a dozen branches.

Step 4: retrieval

Candidates are pulled from a live search index. Which index depends on the platform. Microsoft Copilot runs on Bing. ChatGPT's retrieval has leaned on Bing, with Google-derived signals clearly influencing results by 2026. Google's AI Overviews and AI Mode use Google.

This is where most Indian professional firms die, and they never find out. A site that has never been submitted to Bing Webmaster Tools is frequently not in Bing's index at all, or is indexed at a fraction of its page count. If your pages are not in the index the platform retrieves from, nothing downstream applies. You are eliminated before any judgment of quality is made.

The fix is unglamorous and free: Bing Webmaster Tools, sitemap submitted, IndexNow configured, robots.txt checked so GPTBot, ClaudeBot, PerplexityBot, Google-Extended and Bingbot are not blocked by a template your web developer copied in 2021.

Step 5: reranking

Retrieved candidates get scored and ordered before the model reads them. Relevance to the sub-query, general domain authority, freshness and corroboration all feed in.

Two findings worth holding onto. SE Ranking's November 2025 research found sites with over 32,000 referring domains are around 3.5x more likely to be cited by ChatGPT than sites with under 200, so classical authority still counts for something. The same research found llms.txt has no measurable effect on citation likelihood, so ignore anyone selling it as a strategy.

Freshness matters here too. Content with a recent, visible date tends to be preferred over undated content saying the same thing.

Step 6: grounding and passage selection

Now the model reads the shortlisted documents and picks specific passages to build the answer from. This step is where good content loses to well-formatted content.

02 — WHAT GETS EXTRACTED Same information. One page gets cited. ANSWER BURIED IN PARAGRAPH FOUR THE ACTUAL ANSWER ↓ NOT SELECTED ANSWER IN THE FIRST TWO SENTENCES SELECTED AS CITATION 44.2% of LLM citations come from the first 30% of a document. — LLM citation pattern research, 2026

Research on LLM citation patterns in 2026 found 44.2% of all citations come from the first 30% of a document's text, with a further 31.1% from the middle section. Answers placed at the end of a page are rarely reached.

What gets selected, in order of reliability:

  • A question-shaped heading followed immediately by a direct one or two sentence answer
  • A specific, checkable fact: a number, a qualification, a named service, a location
  • A short list where each item stands alone
  • A table with clear headers

What does not get selected: brochure prose, mission statements, anything that builds to a point over four paragraphs, and anything hedged so heavily it states nothing.

Step 7: synthesis and naming

Finally the model writes the answer and decides whether to name businesses at all. Platforms behave very differently here.

Platform Cites sources Names brands
ChatGPT 87% of answers 20.7% of answers
Google AI Mode 76.3% 37.6%
Google AI Overviews 84.9% 61%

Source: Growth Memo, April 2026

The consequence is that being used as a source and being recommended by name are separate outcomes, and a firm can be doing well at one while failing at the other. ChatGPT in particular behaves like an academic paper: heavy on footnotes, light on naming. If you measure only brand mentions, you will underestimate progress on ChatGPT and overestimate it on AI Overviews.

Why two firms with identical websites get different answers

Because most of the decision is not made on the website.

At step five, corroboration carries real weight. A claim your site makes about itself is one data point. The same claim appearing on your Google Business Profile, your LinkedIn page, the ICAI member directory, JustDial and a review platform is five agreeing data points, and the system can resolve the entity with confidence.

03 — ENTITY CORROBORATION One description, repeated. Or five, contradicting. THE FIRM your website Google Business Profile LinkedIn ICAI directory review platform JustDial old address Sulekha different firm name Reddit thread confirmed unresolved — two inconsistent listings are enough to make the system hesitate

Two things follow.

The first is that inconsistency is not neutral, it is negative. An old address on JustDial and a shortened firm name on Sulekha do not simply fail to help. They introduce doubt about which entity is which, and a system that cannot resolve you confidently will select a competitor it can.

The second is where the corroborating surfaces actually are. SEMrush analysed 325,000 prompts in March 2026 and found LinkedIn is the most-cited domain for professional queries, appearing in 14.3% of ChatGPT Search responses and 13.5% of Google AI Mode responses. Reddit accounts for roughly 40% of AI citations across ChatGPT, Gemini and Claude combined (5WPR AI Platform Citation Index 2026), and 24% of Perplexity citations in January 2026 (Tinuiti). SE Ranking found domains with substantial community presence carry around 4x the citation likelihood, and those with review-platform profiles around 3x.

What you control, and what you do not

Step What decides it Your control Time to move
01 Prompt interpretation The user's wording None n/a
02 Parametric recall Years of public writing about you Very low, very slow Years
03 Query fan-out The platform's rewriting logic None directly; coverage is the response Weeks
04 Retrieval Index presence and crawlability High Days
05 Reranking Authority, freshness, corroboration Medium Months
06 Grounding Content format and specificity High Weeks
07 Synthesis Platform naming behaviour None n/a

Three of the seven steps are entirely outside your influence. That is worth saying plainly, because a provider who implies otherwise is overselling.

The two you control most directly, retrieval and grounding, are also the two that move fastest. That is a fortunate accident, and it explains why a properly sequenced engagement front-loads indexation and content format before touching anything else.

04 — LEVERAGE AGAINST TIME Do the top-left first. Always. WEEK 0 WEEK 24 LOW HIGH LEVERAGE Bing indexation + IndexNow schema markup Google Business Profile answer-format rewrite entity consistency review velocity Reddit and Quora earned placements Indicative, based on delivery across Mumbai professional services engagements.

The four reasons a firm gets left out

Diagnosed in order, because each one makes the next irrelevant.

  1. Not in the index. Eliminated at step four. Nothing else matters yet.
  2. In the index but not resolvable. Described differently across surfaces, so step five cannot establish which entity you are.
  3. Resolvable but not extractable. Step six finds no passage that answers the sub-query directly.
  4. Extractable but not corroborated. Your site says it, nothing else does, so a competitor with five agreeing sources gets named instead.

Most firms try to fix number four with content marketing while failing at number one. That is why their spending produces nothing.

Frequently asked questions

How does ChatGPT decide which businesses to recommend?

ChatGPT rewrites the question into multiple sub-queries, retrieves candidate pages from a live search index, reranks them on relevance and authority, then selects specific passages to build its answer from. A business that is not in the retrieved index is eliminated before quality is ever assessed.

Why does ChatGPT recommend my competitor instead of my firm?

Usually because your competitor is retrievable and you are not, or because their description is consistent across enough third-party surfaces for the system to identify them confidently while yours is not. It is rarely a judgment about which firm is better.

Can you pay to appear in AI search results?

No, not in the organic answer. AI recommendations are produced by retrieval and synthesis, not by placement, and no agency or platform sells inclusion in an organic AI answer. Advertising products inside AI interfaces are separate and clearly marked.

Where to start

Check step four before anything else. Go to Bing Webmaster Tools, add your site, and compare the pages indexed against the pages you actually have. If the number is zero or a small fraction, you have found your problem, and it is free to fix.

After that, open ChatGPT and Perplexity and ask what a prospective client would ask. Ask it five different ways. Whoever appears is surviving all seven steps. You now know exactly which ones they are surviving that you are not.

Lumicite runs a fixed-scope AI Visibility Audit for ₹9,500 that does this properly: 25 test queries across five platforms, scored on a fixed rubric, plus indexation and crawlability checks, competitor intelligence on your top three, and the list of domains AI is currently citing in your category. Five to seven working days, creditable against a full engagement.

Sources

  1. Growth Memo, citation and brand-mention rates across the major AI engines (2026)
  2. SE Ranking, referring-domain thresholds for ChatGPT citation likelihood, llms.txt impact, and community and review-platform presence (2025)
  3. SEMrush, analysis of 325,000 prompts identifying the most-cited domains for professional queries (2026)
  4. 5WPR AI Platform Citation Index, Reddit's share of AI citations across ChatGPT, Gemini and Claude (2026)
  5. Tinuiti Q1 AI Citation Trends Report, Perplexity citation sources (2026)
  6. LLM citation pattern research, distribution of citations across a document's length (2026)
back to blog