evidence-manifest
A small structured record of what produced a benchmark number: eleven fields, a JSON Schema, two checkers, and a reference analyser that turns a directory of these records into a noise floor. Code and data on GitHub.
It came out of The Instrument Moves. Every required field points at a case in that audit where a published score could not be read, because something that shaped it was not recorded.
What two single runs have to differ by
The analyser, run on the repeated runs that each benchmark publishes. The number is the gap two single runs of the same configuration must exceed before it is outside what repetition alone produces (95%).
| Benchmark | Repeated runs | Tasks | Gap (points) |
|---|---|---|---|
| MobileWorld, GUI-only | our audit, 10 rounds | 117 | 8.1 |
| τ²-bench, airline | 4 models × 4 trials | 50 | 10–11 |
| τ²-bench, retail | 4 models × 4 trials | 114 | 7–8 |
| τ-bench v1 | 2 models, 4–8 trials | 50 / 115 | 13 / 8 |
| Terminal-Bench 2.0 | 95 leaderboard configurations | 89 | 7.6 median |
| MCPMark v1 | 67 configurations × 4 trials | 21–30 | 11.5 median |
| METR time horizon 1.1 | 20 agents | 228 | 3.2 median |
| SWE-bench Verified | 12 configurations × 10 runs | 500 | 2.7–4.1 |
On suites of 50 to 115 tasks, a single-run gap under about 7 to 13 points is inside noise. Most leaderboards report single runs. Each row’s table, caveats and source are in the repository’s results/ directory.
Using it
If your benchmark or harness publishes repeated runs, an adapter turns them into manifests and the analyser gives the same table for your suite. There are adapters for MobileWorld, τ²-bench, τ-bench v1, Terminal-Bench 2.0, MCPMark, METR’s time-horizon runs and the ASSERT-KTH artefacts. If you would use it for another harness, say which one: open an issue or write to [email protected].