What we measure, how, and what we never promise.
What we measure
Every audit and every monthly re-audit runs the same battery: 20 localized questions — four intents (discovery, recommendation, comparison, and high-intent tasks like "who does DSCR loans in [city]") — against six AI systems: GPT-, Claude-, Gemini-, DeepSeek-, and Qwen-class chat models plus Google's answer layer. That's 100+ sampled answers per run.
Each response is stored raw, with a UTC timestamp. Reports quote answers verbatim. We never edit, paraphrase, or fabricate a model output — if a report ever shows simulated data (for design previews), it is stamped SAMPLE — SIMULATED DATA on every page.
On top of the raw battery, three analysis passes run on every audit:
- Entity extraction. An extraction model reads every answer and tallies every business named — so share of voice reflects who AI actually recommends, not just the competitors you told us about.
- Mention grading. Every answer that names you is graded against your known facts: sentiment, description accuracy, the exact descriptors used, and any factual errors worth correcting.
- Site & source checks. A live crawl of your website's AI-readiness (does robots.txt block GPTBot/ClaudeBot/PerplexityBot; is there LocalBusiness schema; llms.txt; consistent name-address-phone) plus an indexed-presence check on every domain the AI answers cited — you versus your top competitor.
The score
The BizWhiz Visibility Score (0–100) weights four components:
- Mention rate — 50%. The share of prompt × model runs where your business is named at all. This is the headline number because it is the closest thing to "does AI recommend you."
- Prominence — 20%. First-mentioned reads differently than buried in a list. We measure where in the answer you appear.
- Sentiment / accuracy — 15%. When AI does describe you, is it right — correct city, correct services, no stale or negative claims?
- Source presence — 15%. Whether your site and profiles appear among the sources the answers cite. If AI's sources don't know you, AI doesn't either.
Alongside the score: a per-intent breakdown (exactly which question types you win and lose), true share of voice (every business the answers named), the citation gap table (the domains the answers drew on, with a presence check for you and your top competitor), and a 30/60/90-day action plan assembled from your audit's actual findings. The gap table is the work order: presence on those exact sources is most of the job.
The fix
The Foundation sprint ($1,500, two weeks) repairs what audits most often find broken: missing or malformed structured data (schema), inconsistent name/address/phone across directories, a thin Google Business Profile, no llms.txt, no pages that answer the questions customers ask AI. The Recommended plan ($997/mo) then does the compounding work: building presence on the cited domains, keeping review velocity healthy, refreshing Q&A content, and re-running the battery monthly so progress is measured, not asserted.
What we never promise
- We do not guarantee rankings, mentions, citations, lead volume, or revenue. AI answers are probabilistic and change with time, phrasing, and model versions.
- We do not claim partnership with or endorsement by OpenAI, Anthropic, Google, or Perplexity. We measure their public products.
- We do not use fake reviews, review gating, private blog networks, or anything that puts your profiles at risk. The work is structured data, real content, real citations, real reviews asked for the right way.
What you can hold us to: a documented methodology (this page), a monthly work log of what we did, and timestamped before/after samples you can check yourself.
Why monthly sampling is honest
A single AI answer is an anecdote. A hundred-plus answers, same questions, same systems, month over month — that's a measurement. Trendlines over snapshots; samples over stories.