About

I work on the reliability of agent evaluation: how much a benchmark score moves when nothing changes, why, and what a leaderboard entry would have to record for the answer to be knowable afterwards.

Three things, in the order they happened:

  1. Measured one instrument. Ran one fixed configuration of a mobile-agent benchmark ten times and found the serving endpoint behind the model had changed mid-experiment, with no field anywhere that would have recorded it.
  2. Built the format and the consumer. A small structured record of what produced a benchmark number, and an analyser that turns a directory of such records into a noise floor, a minimum detectable difference, and a drift screen.
  3. Read other leaderboards against their own noise. Where benchmarks publish repeated runs, the analyser runs on them unchanged.

Contact: [email protected] · GitHub