Prompt changes ship on vibes because a single eval pass over a non-deterministic system produces a number, not evidence. Verdict samples every case k times, pairs the runs case by case, and reports a confidence interval — so a change is called a regression only when re-running the same prompt twice wouldn't have produced the same gap.
The agent under test below assesses influencer-marketing creator fit against a brand brief. Pick two prompt revisions and run the comparison live.
| Case | Baseline | Candidate | Δ | Worst grader detail (candidate) |
|---|
Treating 20 cases × 10 samples as 200 independent observations is the fastest way to report an interval several times too narrow. Verdict resamples cases first, then samples within them.
Cases differ enormously in difficulty, and that difference has nothing to do with the change under test. Comparing runs case by case cancels it out.
Every report resamples the baseline against itself and counts how many cases look worse anyway. That number is printed next to the real flag count, so a wall of red can be read honestly.
With too few cases a case-level bootstrap doesn't widen — it produces a falsely tight interval, because it never observed the variance it's missing. Verdict withholds the verdict instead.
Real: the statistics, the graders, the queue, the budget enforcement, the database, and every number on this page — computed live when you press Compare.
Synthetic: the 20-case dataset is hand-written, not scraped, and no creator on it is a real account. The model behind the agent is a deterministic simulator, not an LLM — because testing a regression detector requires a system whose true quality is known, which no real model will tell you. Two of the three prompt revisions differ in true quality; one does not. Verdict is not told which is which.
A live provider implements the same interface and is swapped in with an API key.