Auditingourownpublishedbenchmarkforreproducibility
Real run · Method and claim audit
Check whether a benchmark we had already published could be reproduced by someone outside the company, and rebuild it when the answer turned out to be no.
What it started from
Sources
- The published version 1 results article, twelve tasks run in February 2026, and its scoring rubric as published at the time.
- Every surface on the site that cited its numbers: the answer pages, the footer, `llms.txt`, both translated locales, and the internal skill that writes new answer pages.
- The repository’s own build and route files, for checking which of the article’s claims were still being repeated elsewhere.
Permissions held
- Read and write on a local worktree, plus the authority to mark a published page superseded and rewrite its description.
- No authority to delete the published article. Retraction here means marking it, not removing it.
- No access to any vendor’s account, and no contact with any vendor during the audit.
Constraints
- Defects had to be stated individually and without spin, in public, before any replacement numbers existed.
- The replacement method had to be published before its first run. A method you only see after the results is a method that could have been shaped to fit them.
- Any task depending on an account that is not free and self-serve was disqualified from the replacement suite.
- The replacement had to contain tasks we expect to lose, named in advance and kept in every mean.
Deliberately excluded
- Re-running version 1 to check it. Its ground truth was the live web of February 2026 and nothing was captured, so there is nothing left to check against.
- Any comparison to GAIA, WebVoyager or Mind2Web. The rebuilt suite produces no number comparable to theirs and does not put results in the same sentence.
- Vendor veto. Vendors are notified before publication and may reply on the record; notice is not approval.
- Publishing any replacement score. This run produced a method and a harness. It produced no results, and the first scored run is not until September 2026.
What it actually did
Read the published article and its rubric against one question only: could someone outside this company obtain the same starting state, run the same prompts, and mark the same results without asking us anything?
No approval gate — read-only or reversible— Read-only audit of an already-public artifact.
Enumerate the defects one at a time, each with the evidence for it taken from the published artifact rather than from recollection, and write no remedies until the list was complete.
Ran unattended, human reviewed the output— The defect list was read in full by a human before any remedy was designed.
Decide what to do with the published numbers. Retract them, keep the article online marked superseded rather than deleting it, and rewrite its description so search and social stop advertising the old score.
Human approved before it ran— A human decision about a live public claim, not an agent’s to make.
Rebuild the suite: twenty tasks across eight categories and three difficulty tiers, each with a time budget, a hard stop, and pass/fail checks whose accuracy points sum to exactly 40 and completeness points to exactly 30.
Human approved before it ran
Replace self-reported autonomy with a score computed from a timestamped intervention log, published in two columns (one that discounts approval prompts and a strict one that does not) so the weighting that favours our own product is never the only number on the page.
Human approved before it ran
Name four tasks as unfavourable to Strawberry, in the suite file, with written reasons, before any run. They stay in every mean and the declaration cannot be revised afterwards.
Human approved before it ran— Declaring losses in advance is a commercial decision, taken by a human.
Build the harness so that no field accepts a typed score: the operator supplies observations and the recorder derives all four criteria, recomputing on write and re-deriving on close, so a hand-edited number is rejected rather than published.
No approval gate — read-only or reversible
Generate deterministic fixtures from a fixed seed, with a drift check, storing dates as offsets from the run date so a fixture cannot quietly change what a task is asking two months later.
No approval gate — read-only or reversible
Sweep every remaining surface still teaching the retracted claim: the site footer, the use-case pages, both translated locales, and the internal skill that would have re-seeded the old scores into every future page.
Human approved before it ran
Publish the full method before the first run, including the conflict-of-interest disclosure and the four declared losses.
Human approved before it ran
What came out
Produced
- A published version 2.0.0 methodology, dated before any run, covering the twenty tasks, the scoring schedules, the intervention cost table, the nine invalidation reasons, the reproducibility rules and the monthly cadence.
- A benchmark harness: a versioned suite file of twenty tasks, a published rubric, a results data model with a recorder that enforces it, and a reporter that turns a month of results into a page and a citation line.
- 42 deterministic fixture artifacts across 11 modules, regenerable from a fixed seed with a drift check.
- The version 1 article kept online and marked superseded, with its description rewritten so the old score stops being advertised by search and social previews.
- Four tasks declared unfavourable to Strawberry in writing before the run, namely a fast three-price lookup, a gated directory search, an irreversible CRM merge, and a long unattended crawl, each with the design decision that costs us the points.
Quality checks performed
- The fixture generator asserts every deliberate trap at build time and refuses to emit if one is lost: the boundary rows either side of a renewal window, the two invoices that are genuinely image-only, the grant budget that overshoots by exactly one fixable amount, and the near-duplicate contacts whose "most complete" and "created first" disagree.
- Three answer keys ship empty on purpose, because their correct answer depends on the world on the run date; pre-filling them would be inventing the answer rather than measuring it.
- Three of six pricing snapshots were captured; two vendors returned 403 and one publishes no pricing page. All three failures are recorded as failures with nothing substituted.
- The validator enforces the point sums per task, so no task can quietly become worth more than another, and the recorder refuses a second attempt unless an invalidated first attempt already exists.
- A run cannot be published without a disclosure block, two scorers, at least one scorer independent of every vendor, and a recorded outcome for every cell, either a score or a stated omission.
- The site-wide claim verifier now treats citing any benchmark score as a hard failure, so the retracted numbers cannot return through a future page.
Time and cost
Not shared
No wall clock was kept, and the work was interleaved with the integration-page run on the same night, so no honest split exists between the two. It consumed no Strawberry credits: it ran in a coding agent against a local checkout, not in a companion. A page whose subject is an unmeasured number that got published anyway is not the place to publish an unmeasured number.
Limits
Failure modes observed
- The audit found the defects; it could not undo them. February’s numbers had already been quoted, linked and crawled, and a retraction reliably reaches fewer readers than the claim did.
- Version 1’s results are unrecoverable rather than merely wrong. Its ground truth was the live web with nothing captured, so nobody, including us, can now check what it measured.
- The self-grading defect had been public for months: the article’s headline finding was that a competitor needed a human in the loop, while autonomy was scored by asking each product to rate its own. Nobody reading it caught that. It surfaced only when one specific question was asked of the method.
- Removing the claim took four passes. The fourth found the internal skill that writes answer pages still listing the retracted scores as citable, upstream of all future page copy, so it would have re-seeded the error into the next batch.
- The replacement is unproven. It has produced no results at all. A method that reads well and has scored nothing is not yet evidence of anything, and its first run is the real test of it.
- Twenty tasks a month is a small sample with a wide confidence interval on any single run. The page says so; anyone quoting a single month’s number will not.
When not to run this
- When you are not prepared to retract. An audit whose only permitted outcome is a defence produces a worse artifact than no audit, because it launders the original claim through the appearance of scrutiny.
- When the benchmark underwrites a live commercial claim you cannot pause. Between finding the defects and publishing the replacement there is a period with no number at all, and that period has to be acceptable before you start.
- When there is no independent scorer available. Two of the rules that make the rebuilt suite better than the old one need a person who is not you: an independent scorer, and third-party adjudication when two scorers disagree by more than three points. Without them you have rewritten the rubric and kept the problem.
- Against someone else’s benchmark as a first exercise. Auditing your own published claim is the version of this with the right incentives, and the only version whose findings you are entitled to act on unilaterally.
Run it yourself
Audit this published benchmark for reproducibility. For every task, answer one question: could someone outside this company obtain the same starting state, run the same prompt, and mark the same result without asking us anything? List every place where the answer is no, with the evidence taken from the published artifact itself.
For the rubric, check three things specifically. Who supplies each number, and does any of it come from the product under test. Whether two honest scorers applying the rubric to the same output could land far apart. And whether the ground truth still exists today, or was the live web on the day of the run.
Do not propose a single fix until the defect list is complete, and do not soften a defect because we are the ones who published it. When the list is done, state for each defect what the replacement method does instead, and what it costs us.
Setup
- Put the published article and its scoring rubric side by side, and audit the rubric against the article’s own headline claims rather than against best practice in the abstract.
- Decide the retraction policy before the findings arrive, mark superseded or delete, so the decision is not made under the pressure of a specific finding.
- Write and publish the replacement method before its first run, including the tasks you expect to lose and why.
- Grep the whole estate for the retracted number afterwards, including translations, footers, and any internal instruction file that teaches future writers what is citable.
Connected apps required
None. This run connected to no third-party app.
Other access needed
- The published artifact and its scoring rubric
- Authority to mark a published page superseded
- A local checkout of the site and its content sources
Proof
| Version 1 defect | What version 2 does instead |
|---|---|
| Autonomy was self-reported. Each product was asked to rate its own, while the article’s headline finding was about a competitor needing a human in the loop | Autonomy is computed from a timestamped intervention log kept by the operator, and published both discounted and strict |
| Three of twelve tasks required private accounts, a Sales Navigator seat and our own Attio and Ashby instances, and testers without them were told to score out of 9 | Every account is free and self-serve, and there is one task set for everyone |
| The judge was an LLM grading prose against no ground truth, so a confident, well-formatted, wrong table scored well | Fixed pass/fail checks marked against answer keys held by the scoring team |
| The rubric ran on discretionary ranges, "minus 5 to 10 points per factual error", while scores were quoted to one decimal place | No ranges. A check is pass or fail unless it publishes its own partial-credit schedule |
| Ground truth was the live web with nothing captured, so February’s results cannot be verified by anyone now | Anything live is snapshotted, fixture dates are stored as offsets, and reference answers are built before any product runs |
| Nothing in it was hard for us: a near-perfect score, a clean sweep of all twelve tasks, no headroom | Four tasks declared unfavourable to us before the run and kept in every mean, plus difficulty tiers and hard stops so products can visibly fail |
| There was no run record: no run id, suite version, raw data, product versions or permanent URL | All of those are preconditions of publishing, enforced by the validator rather than by good intentions |
| A GAIA percentage was published with no submission link and no statement of which split it referred to | Not repeated. This suite produces no number comparable to GAIA, WebVoyager or Mind2Web |
A real run is one job, start to finish, with its inputs, its steps and its limits attached. See every published run at real runs.