AI Answer Assurance
Measurement Methodology
How AITWIRE measures, calibrates, and proves the way AI systems answer — what the scores mean, how we calculate confidence, and what the numbers can and cannot tell you. We would rather you trust fewer numbers more.
Everything below is one repeating cycle we call the AI Answer Assurance Loop: probe → judge (agreement‑validated) → fix → re‑measure → prove. Each pass produces the sample sizes, intervals, and controls described on this page.
The AI Answer Assurance Loop
The loop is how a claim earns the right to be reported. Each stage feeds the next, and prove feeds the next probe — measurement never stops:
- Probe. We ask the major AI answer engines controlled questions — run "cold," and repeated across phrasings, engines, and cycles.
- Judge. Each response is scored against facts you've confirmed, and the scorer itself is checked against human-confirmed ground truth (agreement-validated).
- Fix. Measured gaps become prioritized recommendations and data-backed campaign briefs. AITWIRE directs and proves; it does not write your prose.
- Re-measure. The same metric, before and after, on an escalating schedule.
- Prove. We report a movement as real only when it clears the statistics (below) — and, where the content is exclusive to a surface we operate, a controlled retrieval test can distinguish genuine retrieval from chance, with false-positive controls.
How we measure
AITWIRE sends controlled questions ("probes") to the major AI answer engines and scores each response across quality dimensions (accuracy, citation, sentiment, quality, recommendation) and signal categories.
- Probes are run "cold" — no personalization, account history, or AITWIRE context is injected — at low randomness, and repeated across phrasings, engines, and cycles.
- This measures the default answer an un-personalized user is likely to get, as a sample over many queries — not a single check.
- Each response is graded, and results are aggregated per dimension with a sample size and an interval (below).
Judged by a panel — and calibrated by humans
Grading is not one model's opinion. Each response is scored by a panel of independent quality-control judges; where they disagree, the result is held back rather than averaged into a false-confident number. This is the calibrate step between measure and prove, and it has two human checkpoints:
- You calibrate the ground truth. Accuracy is graded against the facts you confirm about your own business — your account's verified canon — so a score measures agreement with what you've verified, not against a guess. AITWIRE suggests; you approve.
- Humans calibrate the judges. On a rolling sample, trained reviewers score responses by hand, and an automated judge is trusted only where it agrees with those humans at a standard inter-rater bar (Cohen's κ). When agreement slips, the affected scores fall back to directional until the judge is re-tuned.
We publish that this calibration happens and the standard measure we hold it to. We do not publish the rubric text, the dimension weightings, or the agreement thresholds themselves — those are what keep the measurement hard to game.
What we measure — and what we don't
AITWIRE measures the engine layer: each AI system is probed through its official, publicly available API, so a score reflects the model's knowledge — plus, where the engine retrieves live, what its retrieval stack returns. The engine layer is the substrate every consumer surface is built on, and it is the layer that can be measured honestly: probes are reproducible (the same question can be re-asked and re-scored), versioned (the resolved model id is recorded with every probe), and run within the engines' published API terms. Consumer applications built on those engines additionally layer personalization, conversation memory, and product behavior that varies per user and per session — a layer that cannot be sampled reproducibly, so we do not claim to measure it.
For the search-grounded engines, the overlap between API and product is direct: ChatGPT Search and Grok's live web search are probed through the same retrieval stacks that power the consumer products, and Perplexity's API is the product.
How faithful each measurement is to the consumer product built on it, across the public AI systems and answer surfaces we cover:
| Fidelity | Engines | What that means |
|---|---|---|
| Strong parity | Perplexity, ChatGPT Search, Grok web search | The API is the product, or shares the product's live retrieval stack. |
| Same model via API | ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google), Grok (xAI), DeepSeek, Mistral | We measure the same model the consumer app runs. The consumer app may add search, personalization, or memory on top. |
| Named proxy | Meta AI | We measure Llama — the model Meta AI is built on — through its API, not the Meta AI consumer product itself. |
| Captured search surface | Google AI Overviews | Not asked as a question — the overview Google actually showed is captured off the results page and scored verbatim. A search surface, not a conversation, so it is reported separately. |
Grok appears twice because it is probed on both routes — its base model and its live web-search variant. It is one engine, not two.
- Google AI Overviews is measured two different ways, and they are not interchangeable. Presence comes from Google Search Console impressions, which Google reports blended with regular search — so that half carries presence data, not per-answer scores. The overview itself is captured off the search results page and scored like any other answer; see how we capture Google AI Overviews below.
- Microsoft Copilot is not probed. Microsoft exposes no public API that returns Copilot's grounded answer, so there is no faithful way to measure it and we do not claim coverage. Where Copilot cites or crawls your content we still track that passively (click-throughs and Bing/bingbot crawler visits) — it never enters an accuracy score.
- Apple Intelligence is not probed, and we do not claim coverage of it.
Where the consumer product itself needs checking, we run UI parity spot-checks: a sample of live probes is re-asked manually, by a human, in fresh logged-out consumer sessions and scored on the same dimensions, so the agreement between the engine layer and the product layer is itself measured. Spot-checks are available per Assurance engagement, and quarterly validation results will be published on this page as they land. The protocol is public: UI Parity Validation Protocol.
Market coverage — what the homepage's 96%+ figure means
The homepage states: "9 systems · 96%+ of the AI-assistant website visits Similarweb tracks worldwide (May 2026)." Here is exactly what that figure is, where it comes from, and — the part that matters most — what it is not a share of.
Source. Similarweb's AI tracker update of 11 June 2026, reporting worldwide all-device traffic to generative-AI assistant websites with a data cutoff of 26 May 2026. That update is the artefact these figures come from; it was the most recent month published when we verified this on 2026-08-04, and it remained the most recent cut the source had published when we re-verified on 2026-08-08. A full transcription of the per-platform figures is available at ppc.land; Similarweb's own narrative reporting on the same panel is here (that page gives the trend in prose, not the per-platform table — so we cite the tracker update, which carries the numbers below).
Denominator. Each platform's share of worldwide all-device traffic to the generative-AI assistant websites Similarweb tracks — seven named platforms (ChatGPT, Gemini, Claude, DeepSeek, Grok, Perplexity, Microsoft Copilot) plus a small unnamed "Other" bucket. The share is a share of visits to that tracked set, and nothing broader.
| System | Share of tracked AI-assistant website visits (Similarweb, May 2026) | Probed by AITWIRE |
|---|---|---|
| ChatGPT | 52.7% | Yes — as two surfaces: ChatGPT and ChatGPT Search |
| Gemini | 27.3% | Yes |
| Claude | 8.9% | Yes |
| DeepSeek | 4.0% | Yes |
| Grok | 2.8% | Yes |
| Perplexity | 1.3% | Yes |
| Microsoft Copilot | 2.0% | No — no public API returns its grounded answer (see above) |
| "Other" (unnamed residual) | ~1.0% | No — not itemised by the source, so excluded from the numerator entirely |
The systems AITWIRE probes sum to 97.0% of that tracked set (52.7 + 27.3 + 8.9 + 4.0 + 2.8 + 1.3). We publish 96%+ — a floor, rounded down — because these are panel estimates that move month to month, and we would rather under-claim than track the decimal. Copilot is the only named tracked platform we do not probe. Two systems we do probe — Meta AI (via Llama) and Mistral — are not in the tracked set at all, so they contribute nothing to the numerator: the figure is conservative on that side too.
What this is NOT a share of — read this before quoting the number. It is not a share of all generative-AI websites worldwide. Similarweb's tracked set is the seven named assistants above; substantial AI sites sit outside it, and some are larger than platforms inside it — character.ai alone drew roughly 182 million visits in May 2026, more than Perplexity or Copilot. Counted against every generative-AI website, the roster's share would fall below the floor we publish. That is precisely why this page and the homepage line both name the tracked set rather than a broader noun, and why we do not describe these as "the seven largest" assistants.
The denominator also excludes native mobile and desktop app usage (apps add materially to every platform's totals), API and enterprise consumption, and embedded assistants (Google AI Overviews inside Search, Meta AI inside Meta's apps, Apple Intelligence). It measures where public assistant usage concentrates on the web — it is not a measure of answer quality, of referral traffic to anyone's site, or of AITWIRE's own performance, and one vendor's panel estimate is directional, not audited.
Refresh. Owner: the AITWIRE account owner. The source publishes monthly, so we check monthly and restate at least quarterly; the figure always carries its measurement month, so staleness is visible on this page rather than buried in a changelog. Headroom above the floor is currently 1.0 point, which is why the check is monthly. If the current data cannot support the stated floor, the number comes down — or comes off — rather than staying up.
How we will capture Google AI Overviews
Status: built, not yet running. This describes a capability that is implemented and disabled. No AI Overview has been captured for any customer, no figures from it appear in any report, and nothing on this page claims otherwise. It is documented here first, deliberately — the disclosure exists before the measurement does, not after. This notice will be removed when capture is live.
Google's AI Overview is the highest-traffic AI answer surface there is, and it is not a chatbot — there is no API to ask it a question. So we will not ask. We will capture the overview Google actually showed on the search results page, and score its text against your canonical facts exactly as we score any other engine's answer. Because the overview names the sources it drew from, citation tracking on Google becomes a record of specific URLs rather than an inference from impression counts.
The boundary, stated plainly. This is a captured search surface, not an assistant conversation. Nobody held a dialogue with Google to produce it. That distinction is why a captured AI Overview is reported as its own class and is never averaged into the assistant-probe figures — blending a search page into a mean of chat answers would make both numbers mean less than either does alone.
- Who captures it. A third-party SERP data vendor, DataForSEO, not AITWIRE. We do not scrape Google. Capturing a search page under our own name would breach Google's terms and put an assurance company in an adversarial position with a measured party — the wrong posture for a business whose product is evidence. The vendor is named here because for an evidence product, who collected an observation is part of the observation.
- Live capture, not cache. Google renders some overviews into the results page and generates others on demand. We always request the on-demand path, and each capture records which of the two produced it — or records that the vendor did not say. "This was observed live" is therefore a property of the record, not an assurance we ask you to take on trust, and where we cannot establish it we store that rather than assume.
- Location and device are pinned and stored. An AI Overview differs by where and on what it is viewed, so each sample carries its location, language and device. Captures will be desktop-pinned; mobile is a genuinely different surface and is not claimed as covered.
- Absence will be recorded, because it is a finding. Google shows an AI Overview for some queries and not others. When none appears, that is stored as a real observation — not dropped, and not counted as a failure — so the share of your questions that trigger an overview at all is computable. A failed capture is excluded from that denominator rather than being counted as "no overview", because a vendor error is not evidence Google withheld anything.
- Volatility — read any single capture as one sample. AI Overviews are re-generated and change between captures, and whether one appears at all can change day to day for the same query. A capture is a sample of what Google showed at one moment, at one location, on one device. It is not the state of the world, and we will not present it as one.
- One known mismatch. Our accuracy judge is written to grade an assistant's reply. An AI Overview is a synthesised page block, so the framing is an imperfect fit. We accept that rather than special-casing the rubric, because changing the rubric would re-baseline every other engine's history to accommodate this one surface.
Two kinds of engine — and why it changes what's fixable
The engines we probe answer in one of two modes, and we report accuracy by engine class because the fix is different for each. (Captured search surfaces, below, are a third class and are reported separately — they are not probes.)
- Grounded engines retrieve live from the web before they answer (for example ChatGPT Search, Perplexity, Grok's web search). What they say responds to the authoritative sources you publish — so a grounded gap is usually movable.
- Training-prior engines answer from what the model already learned in training, with no live lookup. A gap there reflects the model's frozen knowledge — it typically shifts only when the model is retrained, so we label it rather than promise a quick fix.
Splitting the score this way keeps the promise honest: it separates what published facts can move now from what only a future model version will. Each class is reported with its own sample size and interval.
Precision — what "high confidence" means
Every dimension score is reported with its sample size (n) and a 95% confidence interval. We label a dimension "high confidence" only when n ≥ 30 and the 95% interval is within ±7 points.
Because the margin narrows with the square root of n, a noisy (near 50/50) metric typically needs roughly 150–200 probes to reach ±7 points — so "high" usually reflects far more than 30 probes. We always show the actual n and interval, not just the label. Confidence here is the statistical precision of the sample — never certainty.
Validity — the limits the interval does not capture
A tight interval tells you the sample is precise. It does not, by itself, tell you the measured value is true. The honest caveats:
- Scoring can err. Response grading uses automated and machine-learning classifiers, which can mis-grade. We mitigate with independent judge cross-checks and inter-rater agreement — a standard measure (Cohen's κ) of how well the scorer matches human-confirmed ground truth — but do not eliminate error. A tight interval around a mis-scored value is "confidently wrong."
- Samples are not fully independent. Repeated, similar prompts to the same model are correlated, so the effective sample is smaller than the raw count and true intervals can be modestly wider than computed. We discount for this correlation (a standard design-effect adjustment) so confidence is earned, not over-counted.
- Scope. We measure un-personalized, point-in-time responses across a selected set of engines — not every personalized answer, every phrasing, every surface, or future model states. AI systems also drift as their models change.
How we keep the numbers honest
Beyond a precise interval, we apply controls so a number is only reported as meaningful when it earns it:
- Statistical power. When a sample is too small to detect a meaningful change, a non-result is reported as "inconclusive," not "flat" — absence of evidence is not evidence of absence.
- Multiple-comparison control. We test several dimensions at once, so we apply a standard false-discovery-rate correction (Benjamini–Hochberg): a movement is "confirmed" only if it survives correction, otherwise it is flagged "exploratory."
- Representativeness. We stratify probes by query intent (branded, category, competitor, local, buyer-intent) and equal-weight the covered strata, so the headline is not dominated by whichever intent was probed most — and we show the gap versus a naive average.
- Citation integrity. A brand mention is not proof. Only a cited source that actually substantiates the answer counts toward the evidence rate; stale, third-party, competitor, or unresolved citations are labelled, not counted.
- Change vs. model drift. AI engines move on their own. We measure and control for model-wide volatility and report your lift net of it — downgrading attribution to inconclusive when the model itself is too volatile.
- Auditability. Every probe stores a reproducible evidence record — the scorer and rubric versions, the resolved model, and the market — so a historical score can be re-checked under today's methodology.
Evidence you can defend
The point of the loop is not a score — it is a record. Behind every number sits an evidence trail you can inspect and export, not a claim you have to take on faith:
- A reproducible record per measurement — the scorer and rubric versions, the resolved model, and the market — so any historical score can be re-checked under today's methodology.
- An append-only, tamper-evident action log. Privileged actions — measurement runs, verdict decisions, publishes, and account changes — are recorded in a hash-chained log whose integrity can be re-verified on demand. (Append-only and tamper-evident, not "immutable": breaks are detectable, and the chain evidences the record.)
- Exportable with AITWIRE AI. AITWIRE AI workspaces can export their full evidence trail as CSV or JSON — every measurement, verdict decision, signal, publish, and account change — for a compliance or audit review.
What we disclose — and what we protect
We publish our framework so you can judge it. We do not publish the implementation that makes it work — both to keep the measurement hard to game and because parts of it are the subject of pending intellectual property.
| We disclose | We protect |
|---|---|
| The loop, its stages, and their order | Our probe question sets and phrasings |
| The statistical methods we use, by name (all standard) | Scoring weightings and per-dimension thresholds |
| Our confidence-label criteria (n ≥ 30, margin ≤ ±7 pts) | Our judge rubrics and prompt text |
| That every measurement leaves an auditable record | The specific method behind our causal retrieval test |
| That judges are panel-scored and human-calibrated (Cohen's κ), and that accuracy is reported by engine class | The κ thresholds, the judge-panel composition, and per-dimension weightings |
How to read the numbers
- Trust the trend and relative comparison more than any single absolute number.
- Use the confidence label and interval to decide which dimensions are reliable enough to act on; treat "insufficient" or "low" dimensions as directional only.
- The strongest evidence is measured lift — the same metric before and after a change — together with citation delivery, i.e. when an engine actually quotes your published source.
What we do not claim
- We do not claim to read every user's personalized answer.
- We do not guarantee that any AI system will adopt, cite, or correctly interpret your information.
- We do not hold any third-party security certification we have not earned; we state a compliance status only when we can substantiate it.
- Analytics are informational and should not be the sole basis for legal, financial, medical, or compliance decisions.
Versioning
This methodology is versioned. When we change how scores are computed, we restate affected figures and note the change, so period-over-period comparisons stay honest. See the AITWIRE Terms of Service for the legal terms that govern measurement and estimates.