Realruns
Every page here is one job an agent finished, published as a receipt rather than a claim: the sources it was given, the steps it actually executed and which of them a human approved, what came out, what it cost, where it failed, and the template to run it again yourself.
A run that cannot fill every one of those blocks is not published. Where a figure was never measured, duration or credits among them, the page says so instead of estimating it. Both runs below are internal engineering work on Strawberry's own repository, not customer results, and each says so on its own page.
The rules these receipts are held to, covering what gets scored, what we refuse to claim, and the consent required before a customer's work appears here, are published in the Agentic Work Index.
Content production at registry scale
- Turning a tool registry into 96 verified integration pages
Publish one accurate page for every native integration the product ships, with a program, not a reviewer’s attention, deciding whether each factual claim was allowed through.
No connected apps · A local checkout of the monorepo · Node 22.6 or newer · No network access and no third-party credentials
Method and claim audit
- Auditing our own published benchmark for reproducibility
Check whether a benchmark we had already published could be reproduced by someone outside the company, and rebuild it when the answer turned out to be no.
No connected apps · The published artifact and its scoring rubric · Authority to mark a published page superseded · A local checkout of the site and its content sources