← All writing

EXPERIMENTS / 7 MIN READ

What I learnt when I ran 5,040 agent trials.

Prompt sensitivity, flaky success, and why the grader deserved its own investigation.

A small run before a much larger one.

I built Agent Trial Runner around a question that a convincing demo cannot answer: how reliably does an agent complete the same task, and does a configuration change make that reliability better?

I went in expecting the model to be the interesting variable. I was mostly wrong about that, wrong about how long the run would take, and — for a while — wrong about what I was measuring at all. Those three mistakes are most of what I learned.

The intended matrix is 2,500 tasks, 12 configurations, and 10 repetitions: 300,000 trials. The documented dry run was smaller—42 generated travel tasks across those 12 configurations, repeated 10 times, for 5,040 trials. It ran on September 10, 2026. This article is about that dry run, not a completed 300,000-trial experiment.

Each trial gets its own booking records and a pinned prompt. The agent operates through tools, within step, time, and token budgets. The grader compares the resulting records with the task’s expected state. An agent saying it completed a booking does not count as completing the booking.

The prompt and the model are a pair.

The experiment combines four Bedrock models with three prompt styles: terse, careful, and plan-first. I added the prompt axis almost as an afterthought — I expected it to produce a few points of noise around a clear model ranking, and I nearly ran with one prompt to save money.

The prompt turned out to matter about as much as the model, and the reason model averages alone would have hidden it is that the two interact.

Dry-run pass rate over scored trials
Model Terse Careful Plan-first
gpt-oss-20b 98.6% 97.4% 98.1%
nova-micro 84.5% 91.7% 89.5%
nova-lite 82.1% 86.2% 86.9%
ministral-3b 37.1% 80.9% 84.0%

For ministral-3b, moving from terse to plan-first changes the observed pass rate by 46.9 percentage points. For gpt-oss-20b, the three observed rates are close. Prompt tuning against only the strongest model would have given me a very different impression of how much the wording mattered.

I would not turn that into a universal claim about either model or prompt style. This is one generated task set, with 42 tasks and repeated observations of those tasks. The practical lesson is narrower: evaluate the configuration you intend to ship. A prompt’s behavior does not transfer automatically across models.

A pass is an event. Reliability is a pattern.

Repeating each task exposes another dimension. Ministral-3b with the terse prompt was inconsistent on 37 of 42 tasks: it passed on some repetitions and failed on others. Even gpt-oss-20b with the terse prompt had six tasks with mixed outcomes.

One successful attempt cannot tell me whether a user can depend on the next one.

An aggregate rate is useful, but it can hide which tasks are unstable. Repetition makes it possible to separate a consistently difficult task from a task whose outcome changes between attempts. Those two patterns suggest different debugging questions and different levels of confidence in a fix.

Ten repetitions are still a limited sample. And repeated trials of the same 42 tasks do not create 5,040 distinct tasks. I treat this run as a diagnostic experiment, not comprehensive evidence of generalization.

The grader was part of the problem.

For a while I believed the models were worse than they are.

The first numbers had ministral failing about a third of the time and everything else looking mediocre, and that was plausible enough that I nearly moved on. What made me look again was one trace. The agent had been told "I'm lena", searched Nashville, correctly picked the $113 Hilton as the cheapest, booked it with the right dates — and been marked failed. The only difference between what it did and what I expected was that it had written the traveller as "Lena".

My grader compared names exactly. Capitalising one is not a mistake, and I had been failing correct agents for it. On that run it was 28 of 33 failures: 85% of everything I was calling a wrong answer was my own case sensitivity.

What makes that worth writing down is not the bug, which is embarrassing and trivial. It is that from the outside it was indistinguishable from the models being weak, and the fix moved every number on the page. I only found it because I opened a failure instead of trusting a rate.

So I added a control I should have had first: an oracle that completes every task through the same tools the agent uses, and has to score 100%. If it does not, the corpus or the grader is broken rather than the model. It would not have caught this particular bug — the oracle passes names through unchanged — but it draws the line I had been missing, between "the agent got this wrong" and "this task cannot be passed".

After the correction, prominent failure categories included missing hotel bookings, unwanted hotel bookings, and missing flight bookings. An unwanted booking is especially revealing: the agent may have done everything requested and then made an extra state change. A grader that checked only for the requested action would miss that.

Protect the measurement before scaling it.

The runner uses a single coordinating process with 48 workers and an adaptive request limiter. Results live in DynamoDB; non-passing traces go to S3. Since Bedrock requests share an account-level rate limit, adding machines would not automatically remove the bottleneck. Central coordination keeps that constraint visible.

Each matrix cell has an identity. Conditional writes prevent duplicate records when a run resumes, and persisted results let a replacement process continue unfinished work. Isolated state prevents one trial’s booking changes from contaminating another trial.

The dry run recorded zero infrastructure errors and five exhausted trials. Exhausted trials were excluded from the published scored pass rates. That denominator matters: a success rate among completed, scored attempts is different from an end-to-end completion rate. I would keep both completion and exhaustion visible in any release decision.

The repository does not yet implement a release gate or statistical comparison across configurations. These are descriptive results; small differences should not be treated as a definitive ranking.

The metric I meant to measure was missing.

The dry run had one job beyond checking the platform held: measure how fast it actually goes, because nobody publishes your account's real Bedrock limit and the whole schedule for the larger run depends on it.

I did not get it. The summary line scrolled off the terminal, and the rows I had carefully made durable did not include a timestamp — so the one number the run existed to produce was the one number I could not recover afterwards. I had thought hard about not losing results and not at all about being able to answer the question later.

Trials now record when they finished, and the report computes sustained trials per second from the stored rows rather than from anything I have to catch as it goes past. Until the next run, my estimate for 300,000 trials rests on arithmetic — roughly four model calls per trial against a 25-per-second limit, so about thirteen hours — and not on anything observed. I had assumed five, by counting workers instead of the rate limit they all queue behind.

That omission is my favorite lesson from the experiment. Durable outputs need to contain enough information to answer the original question. A completed run is not the same thing as a completed measurement.

Before expanding the matrix, I want the evidence to support the questions the larger run is meant to answer: measured throughput, stable per-task behavior, and comparisons that account for uncertainty. Running more trials only helps when the system preserves what those trials mean.

KEEP READING

A production failure should leave behind a test. ↗