Why the instrument comes first
A method published after its results is a method that could have been shaped to fit them
Almost nothing claimed in this category can be checked by the person reading it. Percentages automated, hours returned, comparative wins: asserted, and rarely reproducible.
We have made the mistake ourselves. Our first competitive benchmark asked each product under test to rate its own autonomy, while the headline finding was about how much human supervision each one needed. Three of its twelve tasks required private accounts nobody outside the company could obtain. Those numbers are retracted.
So the order here is deliberate: publish what will be measured and how, publish what we will not claim, publish the receipts that already exist. Then run it, and publish what it says, including the parts that go against us.
What is measured
Four questions, asked the same way every time
The instrument is small on purpose. Every task in every category is scored on the same four, so a number from one run means the same thing as a number from the next.
Version 2.0.0 of the suite is twenty tasks across eight categories of ordinary browser work (research across tabs, extraction behind a login, form completion, CRM updates, inbox triage, calendar reconciliation, list building and competitive monitoring) in three tiers, from a four-minute lookup to a half-hour grind. Four of the twenty were declared unfavourable to us before any run and stay in every mean.
Did the work finish?
Not whether the agent produced text about the task. Whether the artifact exists: rows in the sheet, the record updated, the form submitted, the brief written.
Was it right?
Scored against an answer key written before the run. A confident wrong answer scores zero, which is the only rule under which reliability becomes legible.
How much human did it take?
Every clarification, nudge, correction and manual step is timestamped and counted against the run. Autonomy asserted by the software under test is not evidence.
Where did it stop?
An agent that halts at an approval gate, a CAPTCHA, or an attestation it cannot honestly make is behaving correctly. Stops are recorded as results, not errors.
Read the full method: every task, its checks, and the scoring rules
What we will not say
The claims we refuse to make, and why
Each of these is a claim we could plausibly make and have decided not to. The reasons are the useful part.
Percentages of work automated
There is no defensible denominator. “90% of admin automated” requires knowing what the other 10% was, and nobody measures that consistently.
Hours or money saved
A before-time nobody recorded is a guess. Where a customer has measured their own baseline and consents, it is their figure to publish, attributed to them.
A score for a product we could not run ourselves
Every product in a comparison is operated by us, on the same fixtures, under the same prompts. A number quoted from someone else’s run, or inferred from a demo, is not a measurement.
Evidence
Receipts are runs someone outside this company could audit
A receipt records the inputs, the steps taken, where a human approved, the result, and the limits observed. The two published so far are our own engineering work and are labelled as such.
One turned a product’s integration registry into ninety-six verified pages behind a gate that checks every factual claim. The other audited our own published benchmark for reproducibility, found eight defects, and replaced it. Both list what went wrong, because a receipt without limits is an advertisement.
How it runs
Fixed tasks, published checks, and numbers nobody types by hand
Tasks are versioned before any run and the suite version is recorded with every result. Fixtures are synthetic and regenerable from a seed, so the starting state can be rebuilt exactly by anyone.
Scoring is pass/fail against published checks. No field in the harness accepts a typed score: every number is re-derived from logged evidence. Two people score independently, and disagreements beyond a threshold go to a third with a recorded note.
Environmental failures are invalidated for one of nine named reasons. A product crashing on its own is a result, not an invalidation. Results are never edited in place; a correction is a new dated version published alongside the original.
Consent
Nothing about a customer’s work appears here without their approval of the finished asset
Not a general permission to be referenced. The specific page, the specific figures, the specific screenshots, approved line by line before publication.
Anonymisation is not a substitute. A metric describing one company’s operations remains theirs to publish whether or not their name is attached. Where consent is absent, the receipt does not exist publicly, and this Index says so rather than describing it vaguely.
Strawberry builds one of the products this programme measures. That is declared on every artifact it produces, including this page.