AI visibility success metrics guide

AI Visibility Metrics: Formulas, Evidence and Audit Examples

Calculate AI visibility metrics with explicit denominators, eligibility rules, worked examples and an illustrative audit. Separate mentions, recommendations and citations.

Decision rule

A useful metric keeps its question set, conditions, and evidence attached.

AI visibility metrics describe observations inside a measurement design. Tracking metrics become success metrics only when they connect to a stated objective. They are not universal market rankings. Before comparing two numbers, confirm that the buyer questions, providers, models, repetitions, country, language, and dates are comparable. If the denominator changes, the story may change even when the underlying answers do not.

Measurement methodology

Define the unit before calculating the percentage.

The unit is one evaluated answer for a query, provider and repetition. Freeze the query inventory, branded/unbranded labels, competitor set, market, language, provider/model and collection window before collecting a baseline. Preserve each raw response and its answer citations separately from retrieval-only sources.

RecoProof filters scoring inputs through its evidence policy: failed, incomplete or uncertain evaluations are excluded, and extraction confidence and brand entity confidence must each reach 0.68. Legacy records without an evaluation status still require both confidence thresholds. This threshold is an implementation rule, not an independently calibrated probability of correctness.

Keep the attempted, eligible and excluded counts together. A missing denominator means not measured; the scorecard omits that metric. A complete eligible answer that does not name the brand is a valid negative observation. A timeout or truncated answer is not. Report delivery is governed separately by the quality gate.

For a first baseline, use the five-question free audit. When a wider sample and human interpretation are needed, use the paid Recommendation Audit. Neither turns a single sample into a market benchmark.

· Eligibility and denominator rules checked against RecoProof scoring and evidence-policy source code; no new external benchmark study.

Illustrative example · Not customer results

From a question to an audit decision.

Five unbranded questions; one hypothetical provider pass; one fixed market and language. No provider was called to create this example. These invented inputs explain the report fields; the layout is a teaching preview, not a screenshot of a completed audit.

Mentioned
2 of 4 eligible answers
Recommended
1 of 4 eligible answers
Needs investigation
1 excluded answer
Example SaaS — four eligible observations and one incomplete answer
QuestionMentionRecommendationOwned citationTracked competitor mentions
Q1: Which tools help a small SaaS team manage onboarding?YesYesYes1
Q2: Which onboarding tools support a self-service workflow?YesNoNo2
Q3: What are alternatives for a team with limited engineering time?NoNoNo1
Q4: Which onboarding tools offer transparent pricing?NoNoNo0
Q5: Which tools fit a security-conscious SaaS buyer?Incomplete answer; excludedNot evaluatedNot evaluatedNot evaluatedNot evaluated
Observation
Q2 names Example SaaS but does not recommend it. Two tracked competitors are mentioned. No owned citation is recorded.
Interpretation
A mention is not a recommendation. Competitor mentions alone do not establish competitor displacement; inspect whether a competitor was actually recommended.
Next action
Open the retained answer and citations, verify the entity classification, and inspect the relevant product page. Retest the same question and conditions after a documented change.
What would the evidence drawer contain?

Exact query, raw response, provider/model, collection time, market/language, response status, entity confidence, and returned citation URLs. There is no real response or citation URL behind this illustrative fixture. An incomplete answer requires investigation; it is not evidence of brand absence, and report delivery remains subject to the quality gate.

Unbranded mention rate

Eligible unbranded answers mentioning the brand / eligible unbranded answers × 100

2 / 4 = 50%

Dividing by five would incorrectly turn 50% into 40%. Keep the excluded answer visible.

Unbranded recommendation rate

Eligible unbranded answers recommending the brand / eligible unbranded answers × 100

1 / 4 = 25%

Q2 is a mention without a recommendation; it does not enter the numerator.

Owned citation rate

Eligible answers with an owned-domain citation / eligible answers × 100

1 / 4 = 25%

Count answers with an owned citation, not the number of links in an answer.

AI share of voice

Customer mentions / (customer mentions + tracked competitor mentions) × 100

2 / 6 = 33.3%

This 33.3% describes the declared sample and competitor set, not market share.

Use the metric definitions and denominator rules to interpret your results. For a wider question set and human review, examine the Recommendation Audit scope.

Metric definitions

What each signal can—and cannot—support.

The final column distinguishes current RecoProof measurement from broader educational concepts. Eligibility and calculation rules were checked against the repository implementation on September 13, 2026; these are product definitions, not independently validated benchmarks.

MetricWhat it measuresWhy it mattersSimple interpretationRecoProof status
Mention rateHow often the brand is named in eligible observations.Establishes basic discoverability before recommendation quality is considered.High mention rate can still coexist with weak recommendations or inaccurate context.Reported from bounded free and paid observations.
Recommendation rateHow often the answer presents the brand as a plausible choice for the stated need.Measures entry into the commercial considered set rather than name recognition alone.Compare only across the same question set and eligible denominator.Reported from free and paid recommendation classifications.
Branded visibilityPresence and accuracy when the question already names the brand.Reveals whether the system recognizes and describes the right entity.A branded win does not establish unbranded category discoverability.Paid scorecard includes branded name presence and answer accuracy.
Unbranded visibilityMentions and recommendations for commercial questions that do not name the brand.Shows whether the product can surface during category and use-case discovery.Usually more commercially revealing than branded recall.Paid scorecard reports unbranded discovery and recommendation rates.
Competitor displacementEligible cases where the target is not recommended and a predefined relevant competitor is.Identifies the exact buyer questions where another vendor enters the shortlist instead.Records the returned set; it does not prove the competitor caused the absence.Classified as query-level evidence in free and paid audits.
Owned citation rateHow often eligible runs cite a page controlled by the measured company.Shows whether owned product, pricing, documentation, or proof pages appear as sources.An owned citation does not automatically mean the brand was recommended.Formal paid scorecard metric; free audit retains owned citation evidence.
Third-party citation visibilityWhich independent sources appear with answers about the category or competitors.Helps investigate external authority and category framing.Source relevance and credibility need qualitative review.Citation evidence is retained; no universal third-party citation score is claimed.
AI share of voiceCustomer mentions as a share of customer plus tracked-competitor mentions in the measured sample.Summarizes relative mention presence inside the declared competitor set.It is sample- and competitor-set-specific, not total market share or a recommendation rate.Current paid scorecard metric with an inspectable numerator and denominator.
Within-model stabilityWhether mention, recommendation, and position repeat consistently for the same query-provider group.Shows whether one observation survives controlled repetition.Stability does not guarantee future permanence.Current paid scorecard metric for repeated groups.
Cross-model agreementWhether provider-level majority mention and recommendation states agree for a query.Separates broad agreement from provider-specific behavior.Agreement is not proof that every consumer interface behaves identically.Current paid scorecard metric for multi-provider questions.
Query and category coverageWhich approved discovery, comparison, alternative, validation, and branded question categories are represented.A rate can look strong because difficult or commercially important questions were omitted.Report the question inventory and category mix beside the outcomes.Recorded in audit question design; not a current named scorecard percentage.
Platform and model coverageThe providers, models, interfaces, countries, and languages included in the measurement.A broader count can improve scope but does not automatically improve evidence quality or comparability.Treat coverage as measurement context, not a success score.Provider and model conditions are recorded; no universal platform-coverage score is claimed.
Visibility trend over timeChange in a like-for-like metric between comparable measurement periods.Shows whether an observed baseline moves after product, evidence, market, or model changes.A trend is invalid when the question set or denominator changes without disclosure.Available through comparable repeat audits or managed monitoring; not inferred from one audit.
Verified RecoProof formulas

How the paid scorecard calculates its core AI visibility KPIs.

These definitions come from the production scoring implementation. Each percentage uses the stated eligible denominator and is rounded to one decimal place. A scorecard metric is omitted when no eligible denominator exists; a missing denominator is not silently presented as a zero-result benchmark.

Unbranded Discovery Visibility
Formula: Unbranded runs with a brand mention ÷ all eligible unbranded runs × 100
Unbranded Recommendation Rate
Formula: Unbranded runs recommending the brand ÷ all eligible unbranded runs × 100
Competitive Consideration
Formula: Eligible competitive runs mentioning or recommending the brand ÷ all eligible competitive runs × 100
Owned Citation Rate
Formula: Runs with a customer-owned domain citation ÷ all eligible runs × 100
AI Share of Voice
Formula: Customer mentions ÷ (customer mentions + tracked-competitor mentions) × 100
Within-model Stability
Formula: Stable repeated query-provider groups ÷ all repeated query-provider groups × 100
Cross-model Agreement
Formula: Multi-provider queries with matching provider states ÷ all eligible multi-provider queries × 100

RecoProof also reports branded name presence, branded answer accuracy, recommendation agreement, citation stability, confirmed inaccuracies, and unverified branded responses when their required samples exist. The composite AI Visibility Index is documented in the RecoProof measurement methodology.

Success by objective

Which AI visibility metrics indicate success?

Which metrics actually matter depends on the decision. Raw mentions answer only “Was the name present?” A commercial measurement should continue through recommendation, citation, and competitor evidence. A brand mentioned in background text but omitted from the shortlist has a different problem from a recommended brand supported only by third-party sources. No single percentage is a universal success benchmark.

Awareness

Use eligible mention coverage to see whether the brand appears at all, without treating presence as commercial selection.

Commercial discovery

Use unbranded recommendation coverage to see whether the product enters the considered set for relevant buyer needs.

Competitive visibility

Use competitor displacement to identify questions where a predefined relevant competitor is recommended while the target is not.

Source authority

Review third-party citation visibility and source relevance to understand which independent evidence supports the category and competitors.

Owned-source strength

Use owned citation coverage to see whether eligible answers surface the company's own product, pricing, documentation, or proof pages.

Longitudinal improvement

Compare equivalent question sets, providers, markets, and conditions over time. RecoProof audits establish a bounded baseline; change requires a comparable retest or monitoring series.

Within-model stability

Does the observation survive repetition?

Repeat the same question under recorded conditions and compare the recommendation, claims, competitors, and citations. Stability does not prove permanence; it indicates whether a finding is robust enough to justify deeper investigation. Report the repetition count and the disagreements instead of hiding them in an average.

Cross-model agreement

Do different systems support the same conclusion?

Compare equivalent questions across specified providers and models. Agreement can increase confidence that an evidence gap is broad; disagreement can reveal provider-specific retrieval or category framing. It should not be converted into a claim that every consumer experience will behave the same way.

Business priority

Start with commercial recommendation gaps backed by inspectable evidence.

  1. 1. Confirm the question matters. Tie it to a persona, intent, and plausible buying decision.
  2. 2. Inspect recommendation and accuracy. Establish whether the brand is considered and represented correctly.
  3. 3. Trace citations and competitor evidence. Identify the source pattern behind the answer.
  4. 4. Check stability. Repeat uncertain findings before allocating work.
  5. 5. Define the action and retest. Use an owner, acceptance criteria, and a recorded measurement window.
See how RecoProof measures this
Query-level examples

Branded and unbranded questions need separate interpretation.

Branded query

“AcmeCRM pricing and integrations”

Measure name presence, description accuracy, owned citations, and unsupported claims. Because the query already supplies the brand, a mention is not evidence of category discovery.

Unbranded commercial queries

“Best CRM for a small SaaS company”

Also test needs such as “Best project management software for agencies,” “[competitor] alternatives,” and “Best analytics tool for startups.” Measure recommendations, competitor displacement, citations, and considered-set coverage.

AcmeCRM is fictional. These are methodology examples, not customer results.

Interpretation also changes by business model. A SaaS team may emphasize unbranded shortlist inclusion and competitor displacement; an ecommerce team may inspect product-category recommendations and the sources supporting current product facts; an agency may need separate, comparable question sets for each client. These examples do not require separate vertical metrics or universal benchmarks.

Measurement questions

AI visibility KPI and benchmark questions.

What is a good AI visibility benchmark?

There is no defensible universal percentage. Start with a declared internal baseline, then compare the same commercial questions, providers, markets, competitors, and classifications over time. A relevant competitor benchmark can add context only when it uses the same sample.

Can I compare AI visibility scores from different tools?

Only after confirming that the tools use comparable prompts, providers or interfaces, locations, dates, repetitions, entity rules, and denominators. Two percentages with the same label can represent different measurements.

How often should AI visibility be measured?

Use a one-time audit when the immediate need is diagnosis. Add a repeated audit or monitoring cadence when launches, competitor movement, or evidence changes can trigger a funded response and someone owns the review.

Measure a baseline

Turn metric definitions into five reviewable observations.

RecoProof's free audit tests five commercial questions and retains the observed recommendations, competitors, citations, and limitations. It is a bounded baseline, not a universal market score.

Run a Free AI Visibility Audit

Build the measurement first

Need the full collection workflow before running a check?

Learn how to perform the audit