A determinism check that passed, and an endpoint that changed anyway

I wanted to measure one number: how much a benchmark score moves when you run the same thing again.

The benchmark is MobileWorld. It drives an Android emulator through a real UI, against self-hosted backends like a mail server and a chat server. Its leaderboard has forty entries. Thirty-seven of them are single runs, and none reports an interval. Entries two or three points apart get cited as evidence that one system is better than another. I wanted to know whether a gap that size means anything.

The plan was simple. Pin everything, repeat the run ten times, report the spread.

Pinning everything

The agent calls a hosted model through an inference router. On a router, one model name can be served by many providers. The model I used, Kimi K2.5, was listed with ten, at several quantizations. Those are different weights on different hardware, so I pinned the provider and the quantization and told the router not to fall back to anything else. I also read the serving provider back from every response, because the benchmark’s own harness had been silently dropping the pin. That turned out to be a bug in how it builds requests, and it has since been reported upstream.

Then the rest: temperature 0, a fixed seed, the container image by digest, the task-set commit, a fresh container for every round, and strictly serial execution so that concurrency could not act as a hidden timing variable.

The check that passed

Early on 25 August (UTC) I was checking that the provider pin actually reached the router. The test request was the prompt Reply with exactly: OK, sent twice to the pinned DeepInfra endpoint at temperature 0. Both responses came back from DeepInfra, which was the point of the test. They were also byte-for-byte identical, 33 tokens each. I wrote down that with pinning in place the output was reproducible.

I had not sent a seed. The answer was 33 tokens long. What I had shown was that one very short, tightly constrained answer came back the same twice. I had recorded it as a property of the endpoint.

The check that failed

Later the same day I tried again with a prompt that looked like the real work: a paragraph asking the model to reason through a phone task step by step before giving an action. Temperature 0, seed 42, reasoning on. Three requests to each of two endpoints, both of which advertise seed support.

DeepInfra, the endpoint that had passed in the morning, returned completions of 2963, 1448 and 1494 tokens. AtlasCloud’s int4 endpoint returned 1571, 2812 and 1427. Six requests, six different answers, with lengths varying about twofold.

So the morning check had given the wrong answer, and in the comfortable direction. I do not think length alone explains it. The two prompts differ in shape as well as length, and on the five endpoints where the daily probe I have run since then sends both kinds of request, short, structured outputs repeat more often than long reasoning ones, but far from always. The lesson I took is narrower: check determinism with requests that look like the workload you care about.

This changed the experiment before it started. I had planned to treat the temperature-0 runs as a measure of environment noise, with the model held still. The model was not held still. So the main run became a measure of everything that is left once every available control is set.

Thirty-three hours

The main run started at 20:10 UTC that evening, pinned to AtlasCloud’s int4 endpoint: 29 tasks, ten rounds, 290 runs, one round after another.

The first three rounds scored 55, 52 and 48 percent. Round four scored 31. Round five scored 31. The remaining rounds landed between 34 and 39, and round six had a separate provider outage that cost it six of its 29 tasks. The score never came back to where it started.

Between rounds one to three and rounds four to ten, ten tasks changed their pass rate. All ten got worse. None got better. Nineteen stayed the same. Repetition noise can move a task either way, so ten out of ten in one direction was the first thing that made me stop.

Looking for my own mistake

My first assumption was that I had broken something. The orchestration was mine and so was the container recycling. But every run is logged with the configuration it ran under, and across all ten rounds there was one image digest, one task-set commit, one model name and one set of pinned parameters. The container was recycled once per round, as intended. Nothing on my side had changed between round three, which ended at 06:47 UTC, and round four, which started at 06:49.

Then I looked at a number I had not been watching: how many tokens the model produced per action. The harness records total tokens per run, but in this workload that is 98.5 percent prompt, and it moved about 8 percent. Completion tokens per action are different. The median was 132.5, 132.4 and 133.1 in the first three rounds, within about half a percent of each other. In round four it was 102.2, then 98.8, 108.7, 105.9, 103.8, 108.4 and 108.0. It fell by about a fifth in one step and stayed there.

The model was producing less reasoning per step, and it was also faster: time per step went from 16.0 seconds to 13.7. A slow network or an overloaded host would make things slower, not faster. Less reasoning per call fits both.

I also checked the time of day. Rounds nine and ten ran overnight, in the same hours as rounds one to three, and stayed at the lower level. A daily load cycle would not do that.

The next day

On 26 August I sent the same long probe again. AtlasCloud, the endpoint the run was pinned to, returned 1344, 1292 and 1292 tokens. The day before, its three answers had ranged from 1427 to 2812. Now two of the three were identical and the third was close. DeepInfra, which I used as a control, still varied: 1938, 1672 and 2427.

So the endpoint I had pinned had changed how it behaved, and the change showed up outside my harness too. I do not know what changed. I can see the behaviour of somebody else’s service, not its configuration, and there is no field on a leaderboard entry that would record it either way.

What the score did and did not show

Looking back, the score did drop. But a round-by-round test on pass and fail counts would not have flagged round four on its own: the p-value was 0.057. The same test on completion tokens per action flagged round four and no other round. That is one event, seen after the fact, so I do not call it a detector. Since September a daily probe has been logging the same kind of signal on fourteen endpoints. Over 328 endpoint-days it raised five single-day length alarms and no alarm that lasted into a second day. That tells me how often it would cry wolf. It cannot tell me how often it would catch a real change, because nothing known changed in that period.

What I would tell someone starting the same experiment

Check determinism with requests that look like your real workload, and send more than two.

Log completion tokens per step. It is one integer per request and it was the clearest trace of the change in my data.

When you report a score from a hosted model, record the provider, the quantization, the date and the image digest. It will not make the number reproducible. It gives the next person a way to notice when two numbers came from different conditions.

The number I set out to measure is in the paper. For this agent on this endpoint, two single runs need to differ by about eight points before the gap exceeds what repetition alone produces. Sixteen of the seventeen gaps between neighbouring entries on the leaderboard are smaller than that. The data and code are on GitHub.