← All writing

EVALUATION DESIGN / 6 MIN READ

A production failure should leave behind a test.

Building evalkeep, and the gap between detecting a mistake and defining the right behavior.

The gap is before the runner.

A production failure gives an engineering team something concrete to fix. It does not automatically give the team a durable way to know whether that fix survives the next prompt change, model update, or tool refactor.

That is the problem I’m working on with evalkeep: turning agent failures into reviewed regression coverage. The project sits upstream of execution. Promptfoo runs the tests; evalkeep helps decide which failures should become tests, what evidence supports them, and what a comparison can actually tell us.

The distinction matters. Running a thousand cases is straightforward compared with knowing whether those thousand cases capture the mistakes users are experiencing. Keeping every trace is not a coverage strategy either. It creates a growing collection with no clear account of what each item protects.

From a trace to something worth keeping.

The pipeline begins with trace ingestion. Inputs are validated and sensitive values are redacted in memory before storage. File adapters support formats including LangSmith exports and OpenTelemetry spans, so bringing in traces does not require handing over vendor credentials.

Failure detection looks for evidence such as explicit failure status, failed evaluators, and negative feedback. Analysis then adds a description of what went wrong. Discovery groups related failures and selects representative cases; dataset generation turns those cases into draft tests.

evalkeep demo .
evalkeep init
evalkeep from-traces refund-agent/traces.jsonl
evalkeep review

This is an intentionally staged workflow. Detection, description, grouping, and test generation answer different questions. Keeping them separate makes it possible to inspect a weak result where it originated, rather than treating the entire pipeline as one opaque transformation.

Knowing what went wrong is only half a specification.

Concretely. A customer asks to refund their latest order; the agent lists the orders and refunds the oldest one. From that trace I can derive, with no judgement at all, that the agent must not call refund_order with order_id = "order-A" — that is the action observed to be wrong, and the tool results the original agent saw travel with the test as fixtures so it replays against the same data.

What I cannot derive is that it should have refunded order-C. Nothing in the trace says so. The recorded conversation contains the mistake and not the intent, and no amount of care in the pipeline changes that.

A trace can show that an agent called the wrong tool. It may not say which tool it should have called, what arguments were required, or what state should have changed. A generated assertion that only forbids the observed mistake is incomplete.

An agent that does nothing can pass a test that only says what not to do.

That is why generated tests remain drafts until review. A person needs to decide whether the assertion captures the failure and whether it also defines useful behavior. Automation can organize evidence and propose coverage; approval is the point where that proposal becomes an engineering commitment.

Once someone has reviewed a test, rerunning the pipeline must preserve that judgment. evalkeep refreshes derived data without overwriting reviews, labels, or edits. Otherwise, the next automated run would erase the most valuable information in the system.

The generated tests need a control, too.

The repository’s tau-bench walkthrough makes this concrete. It starts with 165 recorded tasks from a baseline model, identifies 90 failures, and groups them into four families. In this example, the suite is built with a test for every failure rather than only the representatives.

On 89 comparable cases, the baseline passes 4.5% and the candidate passes 42.7%: a difference of 38.2 percentage points. These are comparisons of recorded trajectories on a failure-selected suite. They are not live production success rates or a general model leaderboard. One case whose target raised an error is excluded from both rates.

The baseline’s low score is useful. The tests came from its failures, so it should fail them. The four cases it passes expose a weakness: without a sufficiently specific failure description, an assertion can target a harmless final lookup instead of the action that caused the problem.

The walkthrough also found a replay mismatch after redaction changed prompt text. A target that missed the lookup returned nothing, and negative-only assertions accepted that empty result. Twelve cases appeared to pass; correcting the mismatch moved the headline by four points. The replay target now raises on a miss.

The public example uses a simulated customer-service benchmark, not production traffic. Its value is that the evidence and replay procedure can be inspected—and that the limitations are visible alongside the score.

Comparison needs the same care. evalkeep separates execution errors from behavioral failures, supports repeated runs so inconsistent cases remain flaky, and uses a paired statistical comparison. When the sample cannot support a confidence interval, it does not manufacture one.

What I want the system to remember.

The useful output is more than a test file. It is a chain from an observed failure to an explicit expectation, a reviewed case, and evidence about whether a later version behaves better.

There is still work ahead: multi-turn replay, tracking failure history over time, and clustering without keeping the full distance matrix in memory. The roadmap keeps those gaps explicit.

The design principle I keep returning to is simple: a regression suite should remember a team’s hard-won understanding of its failures. If the suite forgets the reason for a test—or rewards an agent for doing nothing—it has preserved the artifact and lost the lesson.

KEEP READING

What I learnt when I ran 5,040 agent trials. ↗