Skip to main content

How We Audit AI Visibility

Every AI visibility audit we deliver follows a documented methodology: a fixed prompt set, five comparable metrics, and a repeatable scoring rubric. This page is the methodology — so you can evaluate how the audit works before you buy it.

The scorecard

Five Metrics, Not One Score

A single number hides what matters. We measure the visibility ladder from retrieval to recommendation, per engine, per prompt.

Retrieval (Visibility)

Whether the brand appears anywhere in the AI-generated answer for a prompt. The first rung of the visibility ladder — if you are not retrieved, nothing else follows.

How it is measured

Scored 0/1 per prompt per engine. Visibility = percentage of prompts where the brand is retrieved.

Citation Share

How often the brand's own pages or profiles are named as the source of information in an answer, versus competitor pages or third-party sites.

How it is measured

For each answer, recorded which domains were cited. Citation share = brand citations ÷ total citations across the prompt set.

Recommendation Rate

How often the AI actively suggests the business as the answer to the questioner — the deepest rung of the ladder and the one closest to revenue.

How it is measured

Scored 0/1 per prompt: did the answer recommend the brand without prompting? Recommendation rate = positive prompts ÷ total prompts.

Position

Where the brand appears relative to competitors when multiple businesses are named — first among three, or sixth among ten.

How it is measured

Ordinal position recorded per prompt; aggregated into an average position and a "named first" percentage.

Sentiment

Whether the mention frames the brand positively (recommended, praised), neutrally (listed, described), or negatively (warnings, complaints, corrections).

How it is measured

Each mention classified positive / neutral / negative. Sentiment mix reported per engine and overall.

The process

Five Steps, Every Engagement

1

Prompt Set Construction

We build 20–50 prompts per market: category queries ("best X in London"), comparison queries ("X vs Y"), and problem queries ("how to choose X"). Prompts mirror how real buyers ask, including follow-up phrasing. The set is fixed for comparability across runs.

2

Baseline Run

Every prompt is run across ChatGPT, Gemini, Claude, and Perplexity (and Google AI Search where available), with both desktop and mobile contexts where relevant. Each answer is captured verbatim for the record.

3

Scoring

Each answer is scored against the five metrics for the brand and for 3–5 competitors. The result is a visibility scorecard per engine: retrieval, citation share, recommendation rate, position, and sentiment.

4

Diagnosis

We trace each gap to its cause: citation gap analysis shows which sources the AI used instead of you; site readiness review checks whether your pages are extractable; entity and schema review checks whether the AI knows who you are.

5

Roadmap & Re-measurement

Fixes are prioritised into a 90-day roadmap — quick wins first, then structural work. The same prompt set is re-run monthly (or at the end of an audit-only engagement) so the scorecard shows movement, not opinion.

Reading the score

The Visibility Levels

Scores are always relative to your market: the rubric is calibrated against the competitors in your benchmark, not against a universal average.

LevelScore bandWhat it means
Invisible0–20Rarely retrieved; competitors dominate category answers.
Present20–40Retrieved in some answers, rarely cited, almost never recommended.
Visible40–60Consistently retrieved and occasionally cited; recommendation gaps remain.
Cited60–80Regularly cited as a source; recommended in a minority of relevant prompts.
Recommended80–100A default recommendation for the category in most engines.
Questions

Methodology FAQ

Why do you use 20–50 prompts instead of tracking thousands?

Comparability beats breadth. A fixed, human-curated prompt set can be re-run identically every month, which is what makes trends meaningful. High-volume tools trade this away — thousands of auto-generated queries cannot be replayed consistently, and their scores are not comparable to our rubric.

How is an AI visibility score calculated?

The score is the weighted composite of the five metrics: retrieval, citation share, recommendation rate, position, and sentiment. Weights are agreed during scoping — a business targeting recommendations weights recommendation rate highest; a content site might weight citation share.

Which engines do you test?

ChatGPT, Google Gemini, Anthropic Claude, and Perplexity by default, plus Google AI Search / AI Overviews where available. We add engines (Copilot, Grok, Meta AI) when a client's market uses them.

Do you use third-party monitoring tools?

Sometimes, but the benchmark itself never depends on them. We run the prompt set directly against the engines and record answers verbatim, so the scorecard stands on its own regardless of which software vendors exist or change their pricing.

Can we run the benchmark ourselves after the audit?

Yes — the prompt set and scoring template are handed over with every engagement. Our AI visibility prompts library is the free starting point, and the benchmarks guide explains how to score and trend the results.

Run the Methodology Yourself — or Have Us Run It

Start free with the prompt library and the benchmarks guide, or book the full audit where the same methodology is applied with diagnosis, fixes, and a 90-day roadmap.