on this page
Executive summary
When someone asks ChatGPT for the best CA firm in Mumbai, the model does not consult a ranked list. It runs a sequence of seven steps, and a firm can be eliminated at any one of them without ever reaching the answer. Most firms fail at step four, retrieval, for a reason that has nothing to do with the quality of their website. This post walks through each step, names what decides the outcome at that step, and separates the levers you control from the ones you do not. The practical conclusion is that inclusion is a sequence of gates, so the order you fix things in matters more than how much you spend.
Why "the AI just knows things" is the wrong model
The most common assumption I hear from partners is that AI systems hold a directory of businesses and surface the good ones. They do not.
A model trained on text has a compressed, lossy memory of what appeared in that text. For Deloitte, that memory is dense. For a 12-partner practice in Fort, it is usually empty. So when the question is specific and local, the system goes and looks something up, in real time, using a search index. What it finds in that moment, and how that content is written, decides who gets named.
Being recommended by an AI is not a reputation outcome. It is a retrieval outcome. Those are different problems with different fixes.
The seven steps behind a single recommendation
Step 1: prompt interpretation
The system reads intent, location and any constraints out of the question. "Best CA firm in Mumbai for GST" carries a service, a city and an implied quality judgment. Nothing here is under your control. It matters only because it sets up the next step.
Step 2: parametric recall
The model checks what it already holds. This is where large, heavily-written-about brands win without any retrieval at all.
For an Indian professional firm, assume this is empty. Parametric presence is built over years of being written about in public, at volume, on sources that end up in training data. It is a genuine long-term goal and a useless short-term one. Anyone selling you a three-month plan to get "into the model's knowledge" is selling you something that does not work that way.
Step 3: query fan-out
The system rewrites the question into multiple sub-queries and searches each one. "Best CA firm in Mumbai for GST" might become GST consultants Mumbai, chartered accountant GST filing BKC, top CA firms Mumbai reviews, GST compliance services near me, and several more.
This is the single most consequential mechanic for how the work gets planned. You are not competing for the query the client typed. You are competing for a family of rewritten queries you never see, and your coverage across that family determines whether you survive to the next step. A page optimised for one phrase covers one branch of a tree with a dozen branches.
Step 4: retrieval
Candidates are pulled from a live search index. Which index depends on the platform. Microsoft Copilot runs on Bing. ChatGPT's retrieval has leaned on Bing, with Google-derived signals clearly influencing results by 2026. Google's AI Overviews and AI Mode use Google.
This is where most Indian professional firms die, and they never find out. A site that has never been submitted to Bing Webmaster Tools is frequently not in Bing's index at all, or is indexed at a fraction of its page count. If your pages are not in the index the platform retrieves from, nothing downstream applies. You are eliminated before any judgment of quality is made.
The fix is unglamorous and free: Bing Webmaster Tools, sitemap submitted, IndexNow configured, robots.txt checked so GPTBot, ClaudeBot, PerplexityBot, Google-Extended and Bingbot are not blocked by a template your web developer copied in 2021.
Step 5: reranking
Retrieved candidates get scored and ordered before the model reads them. Relevance to the sub-query, general domain authority, freshness and corroboration all feed in.
Two findings worth holding onto. SE Ranking's November 2025 research found sites with over 32,000 referring domains are around 3.5x more likely to be cited by ChatGPT than sites with under 200, so classical authority still counts for something. The same research found llms.txt has no measurable effect on citation likelihood, so ignore anyone selling it as a strategy.
Freshness matters here too. Content with a recent, visible date tends to be preferred over undated content saying the same thing.
Step 6: grounding and passage selection
Now the model reads the shortlisted documents and picks specific passages to build the answer from. This step is where good content loses to well-formatted content.
Research on LLM citation patterns in 2026 found 44.2% of all citations come from the first 30% of a document's text, with a further 31.1% from the middle section. Answers placed at the end of a page are rarely reached.
What gets selected, in order of reliability:
- A question-shaped heading followed immediately by a direct one or two sentence answer
- A specific, checkable fact: a number, a qualification, a named service, a location
- A short list where each item stands alone
- A table with clear headers
What does not get selected: brochure prose, mission statements, anything that builds to a point over four paragraphs, and anything hedged so heavily it states nothing.
Step 7: synthesis and naming
Finally the model writes the answer and decides whether to name businesses at all. Platforms behave very differently here.
| Platform | Cites sources | Names brands |
|---|---|---|
| ChatGPT | 87% of answers | 20.7% of answers |
| Google AI Mode | 76.3% | 37.6% |
| Google AI Overviews | 84.9% | 61% |
Source: Growth Memo, April 2026
The consequence is that being used as a source and being recommended by name are separate outcomes, and a firm can be doing well at one while failing at the other. ChatGPT in particular behaves like an academic paper: heavy on footnotes, light on naming. If you measure only brand mentions, you will underestimate progress on ChatGPT and overestimate it on AI Overviews.
Why two firms with identical websites get different answers
Because most of the decision is not made on the website.
At step five, corroboration carries real weight. A claim your site makes about itself is one data point. The same claim appearing on your Google Business Profile, your LinkedIn page, the ICAI member directory, JustDial and a review platform is five agreeing data points, and the system can resolve the entity with confidence.
Two things follow.
The first is that inconsistency is not neutral, it is negative. An old address on JustDial and a shortened firm name on Sulekha do not simply fail to help. They introduce doubt about which entity is which, and a system that cannot resolve you confidently will select a competitor it can.
The second is where the corroborating surfaces actually are. SEMrush analysed 325,000 prompts in March 2026 and found LinkedIn is the most-cited domain for professional queries, appearing in 14.3% of ChatGPT Search responses and 13.5% of Google AI Mode responses. Reddit accounts for roughly 40% of AI citations across ChatGPT, Gemini and Claude combined (5WPR AI Platform Citation Index 2026), and 24% of Perplexity citations in January 2026 (Tinuiti). SE Ranking found domains with substantial community presence carry around 4x the citation likelihood, and those with review-platform profiles around 3x.
What you control, and what you do not
| Step | What decides it | Your control | Time to move |
|---|---|---|---|
| 01 Prompt interpretation | The user's wording | None | n/a |
| 02 Parametric recall | Years of public writing about you | Very low, very slow | Years |
| 03 Query fan-out | The platform's rewriting logic | None directly; coverage is the response | Weeks |
| 04 Retrieval | Index presence and crawlability | High | Days |
| 05 Reranking | Authority, freshness, corroboration | Medium | Months |
| 06 Grounding | Content format and specificity | High | Weeks |
| 07 Synthesis | Platform naming behaviour | None | n/a |
Three of the seven steps are entirely outside your influence. That is worth saying plainly, because a provider who implies otherwise is overselling.
The two you control most directly, retrieval and grounding, are also the two that move fastest. That is a fortunate accident, and it explains why a properly sequenced engagement front-loads indexation and content format before touching anything else.
The four reasons a firm gets left out
Diagnosed in order, because each one makes the next irrelevant.
- Not in the index. Eliminated at step four. Nothing else matters yet.
- In the index but not resolvable. Described differently across surfaces, so step five cannot establish which entity you are.
- Resolvable but not extractable. Step six finds no passage that answers the sub-query directly.
- Extractable but not corroborated. Your site says it, nothing else does, so a competitor with five agreeing sources gets named instead.
Most firms try to fix number four with content marketing while failing at number one. That is why their spending produces nothing.
Frequently asked questions
How does ChatGPT decide which businesses to recommend?
ChatGPT rewrites the question into multiple sub-queries, retrieves candidate pages from a live search index, reranks them on relevance and authority, then selects specific passages to build its answer from. A business that is not in the retrieved index is eliminated before quality is ever assessed.
Why does ChatGPT recommend my competitor instead of my firm?
Usually because your competitor is retrievable and you are not, or because their description is consistent across enough third-party surfaces for the system to identify them confidently while yours is not. It is rarely a judgment about which firm is better.
Can you pay to appear in AI search results?
No, not in the organic answer. AI recommendations are produced by retrieval and synthesis, not by placement, and no agency or platform sells inclusion in an organic AI answer. Advertising products inside AI interfaces are separate and clearly marked.
Where to start
Check step four before anything else. Go to Bing Webmaster Tools, add your site, and compare the pages indexed against the pages you actually have. If the number is zero or a small fraction, you have found your problem, and it is free to fix.
After that, open ChatGPT and Perplexity and ask what a prospective client would ask. Ask it five different ways. Whoever appears is surviving all seven steps. You now know exactly which ones they are surviving that you are not.
Lumicite runs a fixed-scope AI Visibility Audit for ₹9,500 that does this properly: 25 test queries across five platforms, scored on a fixed rubric, plus indexation and crawlability checks, competitor intelligence on your top three, and the list of domains AI is currently citing in your category. Five to seven working days, creditable against a full engagement.
Sources
- Growth Memo, citation and brand-mention rates across the major AI engines (2026)
- SE Ranking, referring-domain thresholds for ChatGPT citation likelihood, llms.txt impact, and community and review-platform presence (2025)
- SEMrush, analysis of 325,000 prompts identifying the most-cited domains for professional queries (2026)
- 5WPR AI Platform Citation Index, Reddit's share of AI citations across ChatGPT, Gemini and Claude (2026)
- Tinuiti Q1 AI Citation Trends Report, Perplexity citation sources (2026)
- LLM citation pattern research, distribution of citations across a document's length (2026)