A useful metric keeps its question set, conditions, and evidence attached.
AI visibility metrics describe observations inside a measurement design. Tracking metrics become success metrics only when they connect to a stated objective. They are not universal market rankings. Before comparing two numbers, confirm that the buyer questions, providers, models, repetitions, country, language, and dates are comparable. If the denominator changes, the story may change even when the underlying answers do not.
Define the unit before calculating the percentage.
The unit is one evaluated answer for a query, provider and repetition. Freeze the query inventory, branded/unbranded labels, competitor set, market, language, provider/model and collection window before collecting a baseline. Preserve each raw response and its answer citations separately from retrieval-only sources.
RecoProof filters scoring inputs through its evidence policy: failed, incomplete or uncertain evaluations are excluded, and extraction confidence and brand entity confidence must each reach 0.68. Legacy records without an evaluation status still require both confidence thresholds. This threshold is an implementation rule, not an independently calibrated probability of correctness.
Keep the attempted, eligible and excluded counts together. A missing denominator means not measured; the scorecard omits that metric. A complete eligible answer that does not name the brand is a valid negative observation. A timeout or truncated answer is not. Report delivery is governed separately by the quality gate.
For a first baseline, use the five-question free audit. When a wider sample and human interpretation are needed, use the paid Recommendation Audit. Neither turns a single sample into a market benchmark.
· Eligibility and denominator rules checked against RecoProof scoring and evidence-policy source code; no new external benchmark study.
From a question to an audit decision.
Five unbranded questions; one hypothetical provider pass; one fixed market and language. No provider was called to create this example. These invented inputs explain the report fields; the layout is a teaching preview, not a screenshot of a completed audit.
- Mentioned
- 2 of 4 eligible answers
- Recommended
- 1 of 4 eligible answers
- Needs investigation
- 1 excluded answer
| Question | Mention | Recommendation | Owned citation | Tracked competitor mentions |
|---|---|---|---|---|
| Q1: Which tools help a small SaaS team manage onboarding? | Yes | Yes | Yes | 1 |
| Q2: Which onboarding tools support a self-service workflow? | Yes | No | No | 2 |
| Q3: What are alternatives for a team with limited engineering time? | No | No | No | 1 |
| Q4: Which onboarding tools offer transparent pricing? | No | No | No | 0 |
| Q5: Which tools fit a security-conscious SaaS buyer?Incomplete answer; excluded | Not evaluated | Not evaluated | Not evaluated | Not evaluated |
- Observation
- Q2 names Example SaaS but does not recommend it. Two tracked competitors are mentioned. No owned citation is recorded.
- Interpretation
- A mention is not a recommendation. Competitor mentions alone do not establish competitor displacement; inspect whether a competitor was actually recommended.
- Next action
- Open the retained answer and citations, verify the entity classification, and inspect the relevant product page. Retest the same question and conditions after a documented change.
What would the evidence drawer contain?
Exact query, raw response, provider/model, collection time, market/language, response status, entity confidence, and returned citation URLs. There is no real response or citation URL behind this illustrative fixture. An incomplete answer requires investigation; it is not evidence of brand absence, and report delivery remains subject to the quality gate.
Unbranded mention rate
Eligible unbranded answers mentioning the brand / eligible unbranded answers × 100
2 / 4 = 50%
Dividing by five would incorrectly turn 50% into 40%. Keep the excluded answer visible.
Unbranded recommendation rate
Eligible unbranded answers recommending the brand / eligible unbranded answers × 100
1 / 4 = 25%
Q2 is a mention without a recommendation; it does not enter the numerator.
Owned citation rate
Eligible answers with an owned-domain citation / eligible answers × 100
1 / 4 = 25%
Count answers with an owned citation, not the number of links in an answer.
AI share of voice
Customer mentions / (customer mentions + tracked competitor mentions) × 100
2 / 6 = 33.3%
This 33.3% describes the declared sample and competitor set, not market share.
Use the metric definitions and denominator rules to interpret your results. For a wider question set and human review, examine the Recommendation Audit scope.
What each signal can—and cannot—support.
The final column distinguishes current RecoProof measurement from broader educational concepts. Eligibility and calculation rules were checked against the repository implementation on September 13, 2026; these are product definitions, not independently validated benchmarks.
| Metric | What it measures | Why it matters | Simple interpretation | RecoProof status |
|---|---|---|---|---|
| Mention rate | How often the brand is named in eligible observations. | Establishes basic discoverability before recommendation quality is considered. | High mention rate can still coexist with weak recommendations or inaccurate context. | Reported from bounded free and paid observations. |
| Recommendation rate | How often the answer presents the brand as a plausible choice for the stated need. | Measures entry into the commercial considered set rather than name recognition alone. | Compare only across the same question set and eligible denominator. | Reported from free and paid recommendation classifications. |
| Branded visibility | Presence and accuracy when the question already names the brand. | Reveals whether the system recognizes and describes the right entity. | A branded win does not establish unbranded category discoverability. | Paid scorecard includes branded name presence and answer accuracy. |
| Unbranded visibility | Mentions and recommendations for commercial questions that do not name the brand. | Shows whether the product can surface during category and use-case discovery. | Usually more commercially revealing than branded recall. | Paid scorecard reports unbranded discovery and recommendation rates. |
| Competitor displacement | Eligible cases where the target is not recommended and a predefined relevant competitor is. | Identifies the exact buyer questions where another vendor enters the shortlist instead. | Records the returned set; it does not prove the competitor caused the absence. | Classified as query-level evidence in free and paid audits. |
| Owned citation rate | How often eligible runs cite a page controlled by the measured company. | Shows whether owned product, pricing, documentation, or proof pages appear as sources. | An owned citation does not automatically mean the brand was recommended. | Formal paid scorecard metric; free audit retains owned citation evidence. |
| Third-party citation visibility | Which independent sources appear with answers about the category or competitors. | Helps investigate external authority and category framing. | Source relevance and credibility need qualitative review. | Citation evidence is retained; no universal third-party citation score is claimed. |
| AI share of voice | Customer mentions as a share of customer plus tracked-competitor mentions in the measured sample. | Summarizes relative mention presence inside the declared competitor set. | It is sample- and competitor-set-specific, not total market share or a recommendation rate. | Current paid scorecard metric with an inspectable numerator and denominator. |
| Within-model stability | Whether mention, recommendation, and position repeat consistently for the same query-provider group. | Shows whether one observation survives controlled repetition. | Stability does not guarantee future permanence. | Current paid scorecard metric for repeated groups. |
| Cross-model agreement | Whether provider-level majority mention and recommendation states agree for a query. | Separates broad agreement from provider-specific behavior. | Agreement is not proof that every consumer interface behaves identically. | Current paid scorecard metric for multi-provider questions. |
| Query and category coverage | Which approved discovery, comparison, alternative, validation, and branded question categories are represented. | A rate can look strong because difficult or commercially important questions were omitted. | Report the question inventory and category mix beside the outcomes. | Recorded in audit question design; not a current named scorecard percentage. |
| Platform and model coverage | The providers, models, interfaces, countries, and languages included in the measurement. | A broader count can improve scope but does not automatically improve evidence quality or comparability. | Treat coverage as measurement context, not a success score. | Provider and model conditions are recorded; no universal platform-coverage score is claimed. |
| Visibility trend over time | Change in a like-for-like metric between comparable measurement periods. | Shows whether an observed baseline moves after product, evidence, market, or model changes. | A trend is invalid when the question set or denominator changes without disclosure. | Available through comparable repeat audits or managed monitoring; not inferred from one audit. |
How the paid scorecard calculates its core AI visibility KPIs.
These definitions come from the production scoring implementation. Each percentage uses the stated eligible denominator and is rounded to one decimal place. A scorecard metric is omitted when no eligible denominator exists; a missing denominator is not silently presented as a zero-result benchmark.
- Unbranded Discovery Visibility
- Formula: Unbranded runs with a brand mention ÷ all eligible unbranded runs × 100
- Unbranded Recommendation Rate
- Formula: Unbranded runs recommending the brand ÷ all eligible unbranded runs × 100
- Competitive Consideration
- Formula: Eligible competitive runs mentioning or recommending the brand ÷ all eligible competitive runs × 100
- Owned Citation Rate
- Formula: Runs with a customer-owned domain citation ÷ all eligible runs × 100
- AI Share of Voice
- Formula: Customer mentions ÷ (customer mentions + tracked-competitor mentions) × 100
- Within-model Stability
- Formula: Stable repeated query-provider groups ÷ all repeated query-provider groups × 100
- Cross-model Agreement
- Formula: Multi-provider queries with matching provider states ÷ all eligible multi-provider queries × 100
RecoProof also reports branded name presence, branded answer accuracy, recommendation agreement, citation stability, confirmed inaccuracies, and unverified branded responses when their required samples exist. The composite AI Visibility Index is documented in the RecoProof measurement methodology.
Which AI visibility metrics indicate success?
Which metrics actually matter depends on the decision. Raw mentions answer only “Was the name present?” A commercial measurement should continue through recommendation, citation, and competitor evidence. A brand mentioned in background text but omitted from the shortlist has a different problem from a recommended brand supported only by third-party sources. No single percentage is a universal success benchmark.
Awareness
Use eligible mention coverage to see whether the brand appears at all, without treating presence as commercial selection.
Commercial discovery
Use unbranded recommendation coverage to see whether the product enters the considered set for relevant buyer needs.
Competitive visibility
Use competitor displacement to identify questions where a predefined relevant competitor is recommended while the target is not.
Source authority
Review third-party citation visibility and source relevance to understand which independent evidence supports the category and competitors.
Owned-source strength
Use owned citation coverage to see whether eligible answers surface the company's own product, pricing, documentation, or proof pages.
Longitudinal improvement
Compare equivalent question sets, providers, markets, and conditions over time. RecoProof audits establish a bounded baseline; change requires a comparable retest or monitoring series.
Does the observation survive repetition?
Repeat the same question under recorded conditions and compare the recommendation, claims, competitors, and citations. Stability does not prove permanence; it indicates whether a finding is robust enough to justify deeper investigation. Report the repetition count and the disagreements instead of hiding them in an average.
Do different systems support the same conclusion?
Compare equivalent questions across specified providers and models. Agreement can increase confidence that an evidence gap is broad; disagreement can reveal provider-specific retrieval or category framing. It should not be converted into a claim that every consumer experience will behave the same way.
Start with commercial recommendation gaps backed by inspectable evidence.
- 1. Confirm the question matters. Tie it to a persona, intent, and plausible buying decision.
- 2. Inspect recommendation and accuracy. Establish whether the brand is considered and represented correctly.
- 3. Trace citations and competitor evidence. Identify the source pattern behind the answer.
- 4. Check stability. Repeat uncertain findings before allocating work.
- 5. Define the action and retest. Use an owner, acceptance criteria, and a recorded measurement window.
Branded and unbranded questions need separate interpretation.
“AcmeCRM pricing and integrations”
Measure name presence, description accuracy, owned citations, and unsupported claims. Because the query already supplies the brand, a mention is not evidence of category discovery.
“Best CRM for a small SaaS company”
Also test needs such as “Best project management software for agencies,” “[competitor] alternatives,” and “Best analytics tool for startups.” Measure recommendations, competitor displacement, citations, and considered-set coverage.
AcmeCRM is fictional. These are methodology examples, not customer results.
Interpretation also changes by business model. A SaaS team may emphasize unbranded shortlist inclusion and competitor displacement; an ecommerce team may inspect product-category recommendations and the sources supporting current product facts; an agency may need separate, comparable question sets for each client. These examples do not require separate vertical metrics or universal benchmarks.
AI visibility KPI and benchmark questions.
What is a good AI visibility benchmark?
There is no defensible universal percentage. Start with a declared internal baseline, then compare the same commercial questions, providers, markets, competitors, and classifications over time. A relevant competitor benchmark can add context only when it uses the same sample.
Can I compare AI visibility scores from different tools?
Only after confirming that the tools use comparable prompts, providers or interfaces, locations, dates, repetitions, entity rules, and denominators. Two percentages with the same label can represent different measurements.
How often should AI visibility be measured?
Use a one-time audit when the immediate need is diagnosis. Add a repeated audit or monitoring cadence when launches, competitor movement, or evidence changes can trigger a funded response and someone owns the review.
Turn metric definitions into five reviewable observations.
RecoProof's free audit tests five commercial questions and retains the observed recommendations, competitors, citations, and limitations. It is a bounded baseline, not a universal market score.
Run a Free AI Visibility AuditBuild the measurement first
Need the full collection workflow before running a check?
Learn how to perform the audit