What AITWIRE measures
AITWIRE measures how nine leading AI engines represent you, across six dimensions: accuracy, citation, sentiment, answer quality, recommendation, and competitive position. Every reading carries a sample size (n) and a 95% confidence interval, so you can see how certain the number is.
How a score is decided
Every AI answer is scored against the facts you have confirmed — your own canon — not against a guess. Objective dimensions such as accuracy and citation are checkable against those facts; judgment dimensions such as sentiment and answer quality are graded by an automated judge. A verdict is decided by the whole confidence interval against a threshold set before measurement, not by the point estimate alone.
How the judge is calibrated
An automated judge is trusted only where it agrees with trained humans. On a stratified sample, humans label answers blind — the machine verdict is hidden until they commit — and agreement is measured with standard inter-rater statistics: Cohen kappa, with Gwet AC1 and Krippendorff alpha as cross-checks. A dimension is certified only when the LOWER bound of the kappa 95% confidence interval clears the bar, on an adequate sample. No model is treated as ground truth: even a more capable model stays a calibrated judge, never an oracle.
What kappa means
Kappa corrects agreement for chance. On the standard Landis and Koch scale, 0.60 is the floor of substantial agreement and 0.80 begins almost perfect. AITWIRE treats 0.60 as a floor, not a target, and always reports the actual value with its interval. A floor is monitoring-grade, not a claim of near-perfect agreement.
Assurance levels: choose your bar
You choose the level of assurance you are aiming for. Each level sets a stricter bar, and the level you pick is stamped on your certificate:
- Monitoring — directional tracking. Measured and disclosed, makes no certified claim.
- Standard — certified at substantial agreement (kappa at least 0.60) on an adequate sample.
- High — a higher bar (kappa at least 0.75) with a larger sample and dual-labelled calibration.
- Very High — near-perfect agreement (kappa at least 0.80) with a large sample and reconciled dual labelling, for high-stakes or regulated use.
Choosing a higher level raises your own bar and the effort behind it. It never changes the verdicts and never lowers the neutrality of the instrument. You choose the assurance level — not the verdict.
What certified requires
A dimension is certified only when both are true: the instrument (the judge) is validated at your level, AND your own measurement is precise enough — enough sample, a decisive interval. Where the instrument is still being calibrated, figures are measured and shown but marked not yet certified, so you always know which is which.
Independence
The judge is calibrated once at the platform level and applied the same way for everyone. Your spend can buy more measurement — tighter intervals, a higher level — but can never move the judge or the verdicts. That separation is what makes the assurance mean something.
Where to go next
- See your current readings and intervals in AI Monitor.
- Set the level you are aiming for in your assurance engagement settings.
- Read How to read the numbers for confidence intervals and sample size.