Methodology

A transparent benchmark, not a universal ranking score.

AI answers are probabilistic. The goal is not to pretend otherwise; it is to measure a defined market question set consistently enough to make better decisions.

01

Commercially focused question sampling

Queries are scored for commercial intent, relevance, strategic value, funnel stage, differentiation potential, testability, and likely business value.

02

Controlled provider conditions

Every question receives a controlled first pass. Paid audits repeat only the five highest-value buying questions and provider disagreements, confidence risks, accuracy risks, or citation conflicts, capped at ten repeated questions.

03

Evidence before inference

Mentions, positions, citations, page evidence, and competitor sources are shown separately from the system’s interpretation of the gap.

04

Confidence governs priority

Provider agreement, adaptive repetitions, entity matching, extraction quality, and result consistency determine confidence. Low-confidence findings cannot become high-priority actions automatically.

Read the worked metric calculations to see how excluded answers affect a denominator. Then start with a free five-question baseline, or inspect the human-reviewed audit scope for a wider diagnosis.

Competitor displacement

Measure who enters the recommendation set when the target brand does not.

For one defined commercial buyer question, RecoProof classifies an observed competitor-displacement case only when the answer is eligible, the target brand is deterministically not recommended, and at least one predefined relevant competitor is recommended. Target absence by itself is not displacement. A competitor mention or citation by itself is not displacement either.

Mention

The brand name appears in the retained answer.

Recommendation

The answer presents the brand as a plausible choice for the buyer's stated need.

First choice

Exactly one recommended brand receives explicit preference or ranking language. Mention order alone does not qualify.

Citation

The answer displays a source associated with a claim. A citation is not automatically a recommendation.

Competitor displacement

The target brand is deterministically not recommended while at least one predefined relevant competitor is recommended in the same eligible answer.

The denominator is eligible, deterministically classified observations.

It is not every prompt ever submitted. A completed answer must be eligible for classification and the target recommendation state must resolve to yes or no. Failed, excluded, invalid, or review-required observations do not silently become wins or losses. Review-required evidence stays unclassified until review resolves the ambiguity.

The result is target-relative. If Vendor A is the focal brand and Vendors B or C are recommended while A is not, that is an observed displacement case for A. It does not mean B or C caused A to disappear. The metric records the returned recommendation set; it does not assign causal influence or revenue attribution.

Mention rate asks, “Are we present?” Displacement asks, “When we are absent from a commercially important recommendation, who enters the considered set instead?” See the broader distinction between AI visibility metrics and denominators, then apply the framework with the practical guide to running an AI visibility audit for your brand.

Result variability

One screenshot shows what happened in one run—not a stable market position.

AI visibility observations can disagree because the measured conditions or the tool’s classification design differ. Repetition improves the basis for judging stability, but it does not remove variability or reveal a single “real ranking.” Compare the measurement design before deciding that one dashboard is correct and another is fraudulent.

Prompt design

Wording, constraints, buyer context, and requested output can change the recommendation set even when two prompts represent the same underlying intent.

Run-to-run sampling

The same recorded request can produce different wording or selections. Repetition can reveal disagreement; it does not eliminate it.

Provider and model

Different providers, models, model versions, and interfaces are separate measurement conditions rather than interchangeable replicas.

Retrieval and source context

Where search or retrieval is available, source availability and retrieval behavior can change the evidence surfaced with an answer.

Execution context

Time, language, country, location, session context, and configuration matter only where the measured surface exposes or uses them, so they must be recorded rather than assumed.

Tool methodology

Prompt sets, repetitions, entity normalization, recommendation rules, ambiguous-answer handling, and aggregation logic can make two dashboards disagree.

Prompt sensitivity

Paraphrases are diagnostic comparisons, not independent demand observations.

Two naturally phrased questions can represent the same buyer intent and still return different recommendation sets. That difference is useful evidence about prompt sensitivity. It does not prove which wording is the “real ranking,” and paraphrases should not be counted as separate search-demand events.

RecoProof rerun stability

Repetition is targeted where disagreement or decision value is highest.

The paid audit gives every approved question a first pass across three providers. It then repeats the five highest-value commercial questions plus questions with provider disagreement, entity or extraction confidence risk, accuracy risk, or owned-citation conflict, capped at ten repeated questions. Retained answers support separate within-model, cross-model, recommendation-agreement, and citation-stability views.

RecoProof’s confidence label is a decision rubric based on sample size, entity matching, extraction quality, and provider agreement. It is not a validated statistical confidence interval. Low-confidence evidence remains low priority or review-required rather than being upgraded by presentation alone.

Buyer checklist

Before comparing tools, compare their measurement design.

  • What exact commercial questions were tested?
  • Which providers, models, interfaces, dates, countries, and languages were used?
  • How many first-pass and repeated observations were collected?
  • What exact evidence qualifies a brand as recommended or first choice?
  • How are predefined competitors and entity aliases normalized?
  • Are mentions, recommendations, first choice, and citations reported separately?
  • How are failed, ambiguous, and review-required answers handled in denominators?
  • Can I inspect the retained answer and displayed source evidence behind a finding?
  • How are results aggregated, and what does the confidence label actually mean?

Once a reviewable baseline and response process exist, the same design questions determine whether recurring monitoring is worth paying for. If branded and unbranded prompts diverge, use the focused diagnosis for search-to-recommendation variability.

These model-level measurements need a separate protocol for Google Search. Learn how to measure Google AI search visibility, then compare Google AI Overview tracking tools by retained answers, citation evidence, and collection conditions. The broader AI visibility tools comparison explains which measurement model each purchase supports.

Visible formula

Discovery and repeatability stay separate.

Unbranded discovery × 25% + unbranded recommendation × 30% + owned citation × 15% + within-model stability × 10% + cross-model agreement × 5% + competitive consideration × 15% − accuracy risk × 15%.

Position weights
#11.00
#20.80
#30.60
#40.40
#5+0.20
Put the method in context

See how the evidence becomes an AI visibility audit.

Review the commercial audit scope, compare the fictional sample output, or start with a bounded five-question check.