evidence-manifest

A small structured record of what produced a benchmark number: eleven fields, a JSON Schema, two checkers, and a reference analyser that turns a directory of these records into a noise floor. Code and data on GitHub.

It came out of The Instrument Moves. Every required field points at a case in that audit where a published score could not be read, because something that shaped it was not recorded.

What two single runs have to differ by

The analyser, run on the repeated runs that each benchmark publishes. The number is the gap two single runs of the same configuration must exceed before it is outside what repetition alone produces (95%).

BenchmarkRepeated runsTasksGap (points)
MobileWorld, GUI-onlyour audit, 10 rounds1178.1
τ²-bench, airline4 models × 4 trials5010–11
τ²-bench, retail4 models × 4 trials1147–8
τ-bench v12 models, 4–8 trials50 / 11513 / 8
Terminal-Bench 2.095 leaderboard configurations897.6 median
MCPMark v167 configurations × 4 trials21–3011.5 median
METR time horizon 1.120 agents2283.2 median
SWE-bench Verified12 configurations × 10 runs5002.7–4.1

On suites of 50 to 115 tasks, a single-run gap under about 7 to 13 points is inside noise. Most leaderboards report single runs. Each row’s table, caveats and source are in the repository’s results/ directory.

Using it

If your benchmark or harness publishes repeated runs, an adapter turns them into manifests and the analyser gives the same table for your suite. There are adapters for MobileWorld, τ²-bench, τ-bench v1, Terminal-Bench 2.0, MCPMark, METR’s time-horizon runs and the ASSERT-KTH artefacts. If you would use it for another harness, say which one: open an issue or write to [email protected].