The Instrument Moves: Unrecorded Serving Changes in Agent Benchmarks

Jingyuan Yi · Poster, NeurIPS 2026 workshop Who Verifies the Agents?

We pinned every control a mobile-agent benchmark exposes: model identifier, serving provider, quantization, temperature=0, seed, container image digest, task-set commit, a fresh container per repetition, serial execution. Then we ran the same agent on the same 29 tasks ten times.

What we found

temperature=0 with a fixed seed did not produce determinism on either endpoint we tried. Our first check said otherwise: a short prompt came back byte-identical twice, on an endpoint that turned out not to be deterministic.

Ten and a half hours in, the pinned endpoint’s behaviour changed. Ten tasks got worse and none got better. Median completion tokens per action fell by a fifth at a single round boundary and stayed there. We observed the endpoint’s behaviour, not its configuration: the shift was specific to that endpoint, absent from a control, and invisible to every field a leaderboard records.

Suite score and completion tokens per action across ten repetition rounds; both step down between rounds 3 and 4 and stay lower
Suite score (top) and completion tokens per action (bottom) per repetition round. Both step between rounds 3 and 4 and stay there. The outage in round 6 is a separate event.

Two single-run scores need to differ by roughly eight points before the gap exceeds what repetition alone produces, for this agent, this endpoint and this reweighted task sample. The exact figure is 8.07; resampling whole rounds gives a 95% interval of 5.18–8.53.

The token signal flagged the shift at the round level; the score did not. The signal costs one integer per request to log. It has worked once, after the fact, so we offer it as a side signal and not as a validated detector.

Per-task pass rate before and after the shift for the ten tasks that moved; all ten got worse
Per-task pass rate before and after the shift, one arrow per task that moved: 10 worse, 0 better, 19 unchanged.

Since the paper

Since 4 September the endpoint sentinel has sent the paper’s determinism probe, unchanged, to the paper’s pinned endpoint and to thirteen others, every day. Two things it has recorded bear directly on the paper:

Over the same weeks, three of the eight providers that listed Kimi K2.5 in early September stopped serving it.

What is released

Every run (314 rows, each with the full configuration it ran under), the analysis that reproduces every number in the paper, a script that checks the paper’s numbers against the data, and notes on what the benchmark’s harness actually does. The daily endpoint sentinel grew out of this paper; its false-alarm record is part of the release.

Corrections

Differences from the version submitted for review. No finding depends on any of them.

Cite

@misc{yi2026instrument,
  title         = {The Instrument Moves: Unrecorded Serving Changes in Agent Benchmarks},
  author        = {Yi, Jingyuan},
  year          = {2026},
  url           = {https://github.com/Universeyi/instrument-moves}
}