Research

The AI Visibility Benchmark: what we will publish, and what it cannot prove

A reproducible protocol for measuring how AI assistants describe brands, with pinned prompts, model snapshots, raw answers, and no invented causality.

Flexsent Labs · 6 September 2026 · 3 min read

An AI visibility number is useful only when someone else can reproduce it.

That sounds obvious. It is not. A dashboard can show a percentage while hiding the prompts, the model version, the number of runs, the search settings, and the rule used to count a mention. Without those details, the number is a story with arithmetic attached.

What The Model Says is preparing a benchmark built around the opposite standard: publish the inputs, the raw answers, the evaluation rule, and the limits alongside every result.

What the benchmark measures

The unit is an answer to a question a real buyer might ask an AI assistant. It is not a keyword ranking and not a count of pages indexed by a crawler. The first measurement will use a versioned prompt matrix covering four shapes:

  • direct recommendations;
  • comparisons with a named competitor;
  • category questions;
  • problem-first questions where the category is never named.

The planned first drop will use four pinned model snapshots: ChatGPT, Claude, Qwen, and DeepSeek. The exact provider identifiers, prompt count, repetition count, locale, and retrieval settings will be published in the run manifest before results are reported. This is a collection plan, not a benchmark result. The comparison slice will use a peer roster frozen before collection; it will not be chosen after looking at the answers.

The last shape matters. A brand can rank for its own category phrase and still be absent when a buyer describes the problem in ordinary language. The difference is one of the reasons a single visibility score is too blunt.

Our existing measurement guide explains the basic arithmetic and the ways a measurement can lie. The benchmark turns those rules into a repeatable release.

What every release will contain

Each drop will publish:

  1. the prompt-matrix version and the complete prompt list;
  2. the model provider and pinned snapshot identifier;
  3. the number of runs for every prompt;
  4. the raw answers in machine-readable formats;
  5. the evaluation criteria for a mention, citation, recommendation, and comparison;
  6. both the integer numerator and denominator behind every reported rate;
  7. a changelog explaining what changed since the previous drop.

Prompts will not be silently edited. If a prompt becomes misleading, it gets a new identifier and a new matrix version. Otherwise a time series can appear stable while its measuring instrument has quietly changed.

The public data page will be the index for these releases. The full limits and the no-causation rule live in the methodology, not in a footnote added after the numbers look interesting.

What the benchmark cannot prove

An increase in mentions does not prove that a particular article, technical change, or campaign caused the increase.

Answers are non-deterministic. Results can vary by account, locale, search setting, retrieval state, and model snapshot. A prompt matrix is also a choice made by the researcher; it can favour the positioning it was designed around.

So the benchmark will report movement, not invented attribution. A result can tell us what the chosen prompt set produced under the recorded conditions. It cannot tell us that one intervention caused the change unless a separate design supports that conclusion.

That distinction is the point of the project. A transparent limitation is more useful than a precise-looking number nobody can audit.

When the first drop is ready

This page is a protocol, not a result. The first dataset will be published only when the matrix, model snapshots, raw answers, and evaluation criteria are all recorded together.

Until then, the honest status is simple: the benchmark is in preparation. We would rather publish that sentence than fill the gap with synthetic numbers.