AdvertisingData & TrackingSEOGEOWebsiteCreativeOrganic SocialAll services →
CasesInsightsAboutCareersContact
NLEN
All insights

Measuring AI visibility: does ChatGPT mention your brand?

Measuring AI visibility: does ChatGPT mention your brand?
Jump to8 sections

Measure how often AI systems mention, recommend and cite your brand compared with competitors, using a scalable data-led method.

Does ChatGPT mention your brand when a potential customer asks for advice? Is it recommended, mentioned neutrally or only used as a source? One manual prompt cannot answer that reliably. Measuring AI visibility requires a fixed question set, multiple AI systems, repeated runs and a consistent data model. TNG scales this approach with automation, allowing dozens or hundreds of questions to be tested at the same time.

Why measuring AI visibility is not a one-prompt check

An AI answer is not a fixed search results page. ChatGPT, Gemini and other systems can mention different brands, sources and arguments at different moments. The wording of a question also shapes the answer. “Which agency can help with AI visibility?” may produce a different shortlist from “Who can make my company visible in ChatGPT?”

A consumer app does not always behave like the API from the same provider. A Dutch benchmark published in September 2026 made that visible: 87 questions across seven AI engines generated 1,298 answers. Language models received the same question several times. Even the ChatGPT app and API did not have exactly the same favourites.

A single result therefore carries little weight. You need a sample large enough to reveal patterns and a method that stays consistent in every run. That is the basis of a mature GEO approach.

Start with questions real customers ask

A good measurement programme starts with questions people ask before buying. “Is brand X good?” mainly tests whether AI can find brand information. “Which providers can help me?” tests whether the brand appears spontaneously in the shortlist.

We therefore divide the question set into several types of intent:

  • General discovery: which solutions or providers fit this problem?
  • Category questions: which brands are strong in a specific service or product group?
  • Comparisons: which providers fit a particular budget, organisation or use case?
  • Sector and regional questions: which providers are relevant to this industry or location?
  • Problem questions: how can a familiar customer problem be solved and who can help?
  • Brand questions: what does AI say about the brand, which strengths are mentioned and which alternatives appear?

The first five groups measure unprompted visibility, so the brand name should not be present. Brand questions form a separate diagnostic layer. This prevents a high score caused by inserting your own brand into every prompt.

Turn prompts into a versioned dataset

Each prompt receives a stable identifier. We store its exact wording, language, intent, sector, region and competitors, together with the model, interface, date and settings. An isolated chat becomes a data point that can be run again later.

If the wording changes every month, you cannot tell whether a result comes from improved visibility or from the new question. The core set therefore stays stable and receives a new version when its content changes.

We run important questions more than once to expose variation. A brand that appears in one out of ten responses is in a different position from one mentioned in eight.

How we scale hundreds of AI questions at once

Manual prompting works for an initial exploration, but does not scale and makes results difficult to compare. We send question sets to models and search experiences through automated workflows. Where APIs support batch processing, many requests become one measurement job. Other checks run in parallel or on a schedule.

Every request receives a unique ID, linking the response to its question, engine, repetition and measurement cycle. We retain the raw response and extract brands, positions, sources and tone in a structured analysis layer.

The results come together in a dashboard. Instead of seeing only the latest score, you can identify which prompt clusters improve, where a competitor gains ground and which sources are increasingly used by AI systems.

The KPIs for LLM and AI visibility

A total mention count is not sufficient. A brand may appear often because it is discussed critically or occurs in irrelevant questions. We combine several signals:

  • Mention rate: the percentage of answers in which the brand appears.
  • First mention rate: how often the brand is the first relevant option.
  • Share of voice: the brand's share of mentions compared with selected competitors.
  • Recommendation rate: how often a mention is genuinely phrased as a recommendation.
  • Citation rate: how often the brand's own website is used as a visible source.
  • Source overlap: which external domains support both your brand and its competitors.
  • Message accuracy: whether services, characteristics and evidence are described correctly.
  • Consistency: how stable the result is across repetitions, engines and measurement dates.

These metrics should always be filterable by intent, sector and model. An average can hide the fact that a brand is visible for informational questions but absent when someone asks for a commercial shortlist.

Combine platform reporting with independent testing

Google now provides Search Console information about visibility within generative AI features. It identifies pages that appear as a result or source, but does not answer every question.

Search Console does not explain why Gemini recommends a competitor or how often ChatGPT mentions your brand in a commercial comparison. We combine platform data with prompt measurements to compare mentions, competitors, order, tone and sources across systems.

Turn measurements into practical improvements

The value is what you do with the percentage. When AI recommends a competitor, we examine recurring arguments and sources. Does your brand lack a clear category page? Are cases difficult for crawlers to interpret? Is a service described inconsistently? Do external sources explain the competitor more clearly?

This creates a prioritised action list. The solution may be stronger comparisons, original data or case evidence. Sometimes technical SEO is the bottleneck because information is difficult to crawl or entities are inconsistent. Third-party mentions are useful only when relevant and verifiable.

After a change, we run the same question set again. This shows whether mentions increase, descriptions improve and new sources appear. AI visibility becomes a measurable optimisation process.

What a strong AI visibility measurement delivers

A credible measurement programme does not promise a permanent number-one position in ChatGPT. Generative answers do not work that way. Its value is that it systematically shows where your brand is and is not included in the digital research journey of potential customers.

A fixed prompt set, multiple engines, repeated runs and automated processing create a useful baseline. Every content, technical or authority improvement can be tested against it. The concept remains the same as asking AI questions manually. The difference is scale: hundreds of controlled questions instead of one screenshot.

Meer lezen