There is no single AI engine that automatically understands every business best. Test ChatGPT, Gemini and Perplexity with the same buyer-intent prompts. Score each answer for correct business identity, services, location, sources, outdated facts and whether the user can reach a useful page. The engine with the most mentions is not necessarily the strongest result if those mentions are inaccurate or commercially irrelevant.

If ChatGPT names your business and Gemini does not, that does not prove ChatGPT “ranks you higher”. The products retrieve, combine and present information differently.

The useful question is narrower:

Which engine gives a prospective customer the most accurate and useful understanding of this business for the prompt being tested?

Build a scorecard around that question.

Use the same prompt in every engine

Comparisons fail when each product receives a different question.

Choose a fixed set of prompts and paste them without rewriting them to help a particular engine.

Test at least four intent types:

  • service discovery;
  • local discovery;
  • provider comparison;
  • trust or proof.

For a Pretoria web design studio, the set could include:

Who designs websites for service businesses in Pretoria?
Compare three web designers that serve Johannesburg businesses.
Which South African web design providers publish clear pricing?
Which website designers show real case studies and lead-generation work?

Run the prompts on the same day where practical. Record the exact wording.

Score identity before visibility

A mention is only useful when the engine identifies the right business.

Check:

  • official business name;
  • website domain;
  • location or service area;
  • core service category;
  • phone or contact route where shown;
  • whether the answer mixes the business with a similarly named company.

Give an identity error more weight than a missing marketing phrase.

If an engine invents a branch or attributes another company’s reviews to you, mark the result as materially inaccurate even if your name appears first.

Test whether the engine understands what you sell

Many businesses describe themselves with broad language that hides the actual service.

An engine may know the company exists but fail to connect it with a buyer’s problem.

Compare the answer with your current offer:

Question Correct?
Does it name the core service?
Does it include the highest-value service?
Does it mention a service you stopped selling?
Does it understand the customer type?
Does it understand the real service area?

If all three engines repeat the same old service, investigate the sources carrying that information before rewriting the homepage.

Record the sources, not only the prose

ChatGPT Search can provide links to web sources. Perplexity is built around cited web answers. Gemini can use Google services and public information, depending on the feature and question.

For each result, record which source appears to support the business description.

Classify it:

  • official website;
  • Google Business Profile or Maps data;
  • customer review;
  • directory;
  • client or partner page;
  • news or editorial article;
  • outdated profile;
  • unknown or unsupported claim.

A correct answer supported by an old directory can still expose a maintenance risk. If that directory changes or disappears, the source mix may change.

Measure commercial usefulness

An answer can be factually correct and still fail the buyer.

Ask whether the result helps someone take the next step.

Useful outcomes include:

  • link to the relevant service page;
  • clear explanation of what the company does;
  • accurate service area;
  • pricing context where available;
  • evidence such as case studies;
  • a working contact route.

Less useful outcomes include:

  • homepage mention with no service context;
  • an outdated social profile;
  • a directory page with the wrong telephone number;
  • a broad blog article when the user asked to hire someone.

Score the landing destination separately from the answer text.

Use an accuracy scale that forces a decision

Avoid a vague 1-to-10 score without definitions.

Use four levels:

0 — Missing
The business does not appear for a relevant prompt.

1 — Misleading
The business appears but important facts are wrong or mixed with another entity.

2 — Accurate but weak
Core facts are correct, but the answer lacks evidence, service detail or a useful route.

3 — Accurate and commercially useful
The answer identifies the business correctly, matches the user’s need and connects to credible evidence or a useful page.

This makes month-to-month comparisons easier.

Do not treat personalised answers as universal rankings

AI products can respond differently based on location, account context, conversation history, product mode and current web results.

Record conditions that could matter:

  • signed in or signed out;
  • location;
  • web search enabled;
  • follow-up context;
  • date;
  • device or product surface when relevant.

Do not publish “We are #1 on ChatGPT” because one account received one answer.

Repeatable observations are more defensible than a ranking claim.

Compare failure patterns

The value of the three-engine test is finding patterns.

All three miss the business

Check foundational discoverability:

  • crawl access;
  • indexability;
  • service-page clarity;
  • business identity consistency;
  • public third-party footprint.

One engine understands the business well

Inspect the sources it uses. You may find a page or third-party profile the others do not surface.

Do not copy the answer. Identify what evidence made the result possible.

All three name the business but get the same fact wrong

Find the source of the wrong fact.

Correct business-controlled pages first, then update or request corrections from relevant third-party sources.

The engines understand the business but recommend competitors

Now the problem is less about entity recognition and more about comparative evidence.

Compare:

  • service depth;
  • case studies;
  • reviews;
  • local evidence;
  • pricing clarity;
  • third-party mentions;
  • topical authority;
  • page quality and conversion path.

Add Google AI as a fourth surface when relevant

The downloadable baseline on IDJoy includes Google AI alongside the three engines in this article.

Google’s AI Overviews, AI Mode and Maps experiences use Google’s search and place information in ways that differ from Gemini’s standalone app. If local discovery matters, test Maps and Google Search surfaces as separate observations rather than assuming Gemini represents all Google AI behaviour.

Run the scorecard monthly

Do not optimise to a single day’s answer.

Keep the same prompt set and compare:

  • mentions;
  • factual accuracy;
  • sources;
  • landing pages;
  • commercial fit;
  • qualified AI-referred enquiries where measurable.

Add a new prompt when the business launches an important service, enters a new genuine service area or discovers a recurring customer question.

Retire prompts that no longer match the offer, but keep the historical result.

Use the comparison to choose work, not a winner

The goal is not to declare ChatGPT, Gemini or Perplexity the “best AI”.

The goal is to identify what each engine currently understands about your business and what evidence is missing.

That gives you a practical work list: fix an old location, open crawler access, improve a service page, publish a case study, correct a directory or make the landing page easier to act on.

IDJoy’s AI visibility audits use cross-engine testing to find those gaps without treating one AI answer as a permanent ranking.