Skip to content
anserra

Methodology

How Anserra measures, and how wrong it can be

AI answers are not deterministic. Any tool that shows you a single number without a range is hiding that. This page explains what Anserra samples, how it calculates intervals, how the score is weighted, and where the limits are.

Updated:

1. Prompt sets

Each business gets a prompt set: the questions a customer would ask an assistant when looking for that kind of business in that place. Prompts are grouped into three kinds, because they behave differently:

  • Category prompts — “best family dentist in Cedar Park”. The business is competing to be named at all.
  • Task prompts — “dentist near me open Saturday that takes new patients”, “how much is a crown in Austin”. Conditions and prices matter.
  • Brand prompts — “Cedar Park Dental reviews”. The engine knows the name; the question is whether what it says is accurate.

The generated set is a starting point. You can edit prompts within your plan's limit (25 per business on paid plans). Brand prompts are kept to a minority of the set so that the aggregate reflects discovery, not recognition.

2. Engines and how answers are collected

Base engines: ChatGPT (OpenAI), Perplexity, Google AI Overviews. Add-ons: Claude (Anthropic), Gemini (Google). Answers are requested through each provider's API with search or browsing enabled where the API offers it, with location context set to the business's city. Anserra does not scrape chat interfaces and does not use logged-in consumer accounts. This means results reflect the model and its retrieval, not a particular user's history — which is what you want for a measurement, and also why an individual user's chat can differ from the report.

3. Samples and confidence intervals

Language models sample their output. Ask the same question three times and you can get three different lists. Anserra therefore sends each prompt to each engine 3 times per run on paid plans (once on the free scan) and treats “named / not named” as a binomial outcome.

For a single prompt with 3 samples the interval is wide: 1 of 3 mentions is 33% with a 95% interval of roughly 6–79% (Wilson score). That is honest and not very useful on its own. The interval narrows as prompts are aggregated: across 25 prompts × 3 samples on one engine, 75 outcomes, a mention rate of 29% carries an interval of about ±10 points. Across three engines, 225 outcomes, about ±6.

Anserra shows the interval next to every rate and reports a change only when the new rate falls outside the previous interval. Week-to-week wobble inside the interval is shown in the chart but not announced as a change.

AggregationOutcomes per runTypical 95% interval at 30%
One prompt, one engine3≈ ±40 points
All prompts, one engine75≈ ±10
All prompts, three engines225≈ ±6
Four weekly runs, three engines900≈ ±3

4. What is recorded per answer

  • Whether the business is named (exact name or unambiguous variant)
  • Position among named businesses
  • Statements about the business, checked against the site: address, hours, phone, services, prices
  • Other businesses named
  • Cited URLs, where the engine returns them
  • Whether the answer contains a path to action: link, phone, booking mention

5. The score

The score is a weighted sum of four components, each 0–100, always shown with its interval and alongside the components:

ComponentWeightWhat goes in
Mention rate40%Share of answers naming the business, category and task prompts weighted above brand prompts
Accuracy15%Share of statements about the business that match the site
Source presence15%Share of citation weight on domains where the business is listed correctly
Agent readiness30%Tasks passed in the agent run, with the site analyzers A1–A6 as a prior when no run has happened yet

Agent readiness has a large weight on purpose. A business that is named but cannot be booked by an agent loses the customer at the last step; the score should say so.

6. The agent run

AnserraBot opens the site in a real browser with a desktop viewport, identifies itself in the user-agent, honors robots.txt, and attempts five tasks by following links already present on pages. It does not type into fields, does not submit forms, and stops when it reaches a booking or checkout URL. Each task ends in Passed or Stopped, with the URL and a screenshot of the stopping point. Levels: L0 no task passed; L1 tasks 1, 4 and 5 passed; L2 additionally tasks 2 and 3; L3 additionally a valid /.well-known/agent-service.json. The specification is open.

7. What the numbers cannot tell you

  • Not traffic. A mention rate is not a click rate. Engines do not report how many users asked the prompt.
  • Not a ranking. Position varies more than mention; treat “best position” as anecdote, mention rate as the measure.
  • Not every user's answer. Consumer apps personalize with memory and location. The API measurement is the cleanest common baseline, not a prediction of one person's chat.
  • Not stable across provider updates. When a provider changes its model or retrieval, rates can shift for everyone. Anserra marks known provider changes on the timeline.
  • Not causation. A rate that rises after a fix is evidence, not proof. Four weekly runs before and after, outside the interval, is the standard Anserra uses before calling a change real.

8. Data retention and research use

Raw answer text is kept for 12 months; aggregates for the subscription plus 30 days. Anonymized, aggregated results from free scans may be used in public research; no business is identified without consent. Details in the Privacy Policy.

See what AI says about your business today

One business, one-time check. No account, no card. Results in about 60 seconds.