VERDICT

regression testing for LLM agents

Your agent got worse. Or it didn't. One run can't tell you.

Prompt changes ship on vibes because a single eval pass over a non-deterministic system produces a number, not evidence. Verdict samples every case k times, pairs the runs case by case, and reports a confidence interval — so a change is called a regression only when re-running the same prompt twice wouldn't have produced the same gap.

The agent under test below assesses influencer-marketing creator fit against a brand brief. Pick two prompt revisions and run the comparison live.

Live comparison

Per-case breakdown

CaseBaseline CandidateΔ Worst grader detail (candidate)

The prompts being compared

Baseline
Candidate

Why this is harder than it looks

Cases are the unit, not samples

Treating 20 cases × 10 samples as 200 independent observations is the fastest way to report an interval several times too narrow. Verdict resamples cases first, then samples within them.

Pairing removes the loudest variance

Cases differ enormously in difficulty, and that difference has nothing to do with the change under test. Comparing runs case by case cancels it out.

The noise floor is measured, not assumed

Every report resamples the baseline against itself and counts how many cases look worse anyway. That number is printed next to the real flag count, so a wall of red can be read honestly.

No verdict below a minimum

With too few cases a case-level bootstrap doesn't widen — it produces a falsely tight interval, because it never observed the variance it's missing. Verdict withholds the verdict instead.

What is real here and what is not

Real: the statistics, the graders, the queue, the budget enforcement, the database, and every number on this page — computed live when you press Compare.

Synthetic: the 20-case dataset is hand-written, not scraped, and no creator on it is a real account. The model behind the agent is a deterministic simulator, not an LLM — because testing a regression detector requires a system whose true quality is known, which no real model will tell you. Two of the three prompt revisions differ in true quality; one does not. Verdict is not told which is which.

A live provider implements the same interface and is swapped in with an API key.