Teams that write evaluations before writing prompts ship fewer production incidents and catch regressions 11x faster — the discipline is eval-driven development, and the gap between teams who practice it and teams who don't is now measurable in both reliability and iteration velocity.
The Mistake Most Teams Make
There is an observable diagnostic for teams that don't have an evaluation discipline: they describe their LLM system's behavior with words like "inconsistent," "sometimes it works," and "I'm not sure why it changed." These are not descriptions of an LLM problem. They are descriptions of a measurement problem.
If you can't specify, before deployment, what correct behavior looks like on a representative sample of inputs, then you have not finished designing the system. You've built something you can't reason about operationally.
Eval-driven development (EDD) is the discipline of writing evaluations before writing prompts, running them as part of CI, and treating a regression in eval scores as a deployment blocker. It is the LLM-system equivalent of test-driven development, with one critical difference: because LLM outputs are non-deterministic, "passing an eval" means "meeting a quality threshold on a distribution of outputs," not "producing an exact string match." The tooling and mental models are different, but the discipline is the same.
The V1 piece in this series documented how switching from string-equality assertions to semantic-equivalence eval harnesses reduced post-deploy incidents to zero for two consecutive quarters. This piece is about the workflow that produces that outcome — the specific steps, tooling, and measurement architecture.
Why It Persists
The dominant pattern in LLM development is not eval-driven — it's vibe-driven. A developer writes a prompt, runs it against a few examples by hand, observes that the outputs "look right," and ships. The eval harness gets built when something breaks in production. At that point, the team is under pressure, the eval cases are written to reproduce the failure rather than to characterize the system's behavior, and the harness is permanently reactive rather than proactive.
This pattern persists because eval harness infrastructure has historically been expensive to build and because there is a psychological trap: it is genuinely hard to write evaluations for behavior you haven't observed failing yet. "What should I test?" feels harder than "test what broke." But this is exactly backwards — the value of an eval harness is highest before you know what failure looks like, because the harness is what tells you when failure is happening.
Hamel Husain's practitioner writing on LLM evals ("Evals are All You Need") articulates a principle I've come to agree with: the best eval cases are not written by engineers, they are written by domain experts observing the system. Engineers write test cases that match their mental model of the system's behavior. Domain experts write cases that expose the gaps in that mental model. A healthcare system should have eval cases written by clinicians. A contract review system should have eval cases written by lawyers. The eval harness is only as good as the cases it runs.
What the Literature Actually Says
The OpenAI Evals framework (2023) established a practical taxonomy for LLM evaluation that the field has converged around: accuracy evals (does the model produce the correct answer?), safety evals (does the model avoid harmful outputs?), consistency evals (does the model produce similar outputs for semantically equivalent inputs?), and calibration evals (when the model is uncertain, does it say so?).
The Anthropic model card methodology extends this taxonomy with behavioral evals — does the model follow the intended policy when the policy is ambiguous or under adversarial pressure? This category matters enormously for production systems and almost no team is running it.
The practical challenge for production teams is that accuracy evals require a ground-truth answer to compare against, and many production tasks don't have ground truth. A summarization task doesn't have a single correct output. A question-answering task over proprietary documents may have multiple acceptable phrasings of the correct answer. This is where LLM-as-judge evaluation becomes the practical tool: use a separate, capable model (often a larger or more capable model than the one being evaluated) to assess whether the output is acceptable against a rubric.
At $0.04 per eval case (using GPT-4o-mini or Claude Haiku as the judge), running 500 eval cases on every pull request costs $20. This is not a budget constraint — it is a rounding error in the cost of deploying a change that breaks a production feature.
The Architecture That Works
The eval-driven development workflow has four phases:
EDD WORKFLOW
=============
PHASE 1: BEHAVIORAL SPECIFICATION
-----------------------------------
Before writing a single prompt:
→
Define the system's intended behavior in 5-10 behavioral assertions
→
Write 20-50 seed eval cases per assertion (human-authored)
→
Define the judge rubric: what criteria determine PASS/FAIL?
→
Store cases in version control alongside the prompt
PHASE 2: PROMPT DEVELOPMENT
------------------------------
Write the prompt against the eval suite:
→ Run
evals locally on the seed set
→
Iterate prompt until seed set passes (target: >85% accuracy, >95% on safety/consistency)
→
Record: which cases fail, and why?
PHASE 3: CI INTEGRATION
--------------------------
On every PR that modifies prompt, retrieval, or model config:
→ Run
full eval suite (seed + accumulated regression cases)
→
Compute delta vs. main branch score
→
Block merge if score regresses >5% on any axis
→ Post
eval report as PR comment
PHASE 4: PRODUCTION MONITORING
---------------------------------
In production, sample 2-5% of real requests:
→ Run
LLM-as-judge on sampled outputs
→
Alert if rolling 24h pass rate drops below threshold
→
Automatically promote failing cases to regression suite
→
Review failing cases weekly to identify new behavioral categories
|
The key architectural decision is where eval cases live and how they are maintained. Eval cases in a spreadsheet that someone maintains manually will not stay synchronized with the system. Eval cases in version-controlled JSON, loaded by the same CI pipeline that builds the system, will.
Here is the eval case schema I use:
{
"id": "contract-review-001", "category": "termination-clause-extraction", "input": {
"document_text": "...", "question": "What is the notice period required for
termination?"
}, "expected_behavior": "The response must identify the
notice period (numeric value + unit), cite the relevant clause section, and
express uncertainty if the clause is ambiguous or absent.", "judge_rubric": {
"contains_numeric_period": { "weight": 0.40, "method": "regex" }, "cites_section": { "weight": 0.30, "method": "llm_judge" }, "handles_ambiguity": { "weight": 0.30, "method": "llm_judge" }
}, "failure_mode_tag": "extraction", "ground_truth": "90 days, Section 8.2", "added_by": "legal-sme", "added_date": "2024-11-15", "regression_source": null
}
|
The judge_rubric field is the critical innovation here: different criteria within the same eval case can use different evaluation methods. Numeric extraction is a regex check — cheap, deterministic, fast. Section citation and ambiguity handling require semantic judgment and use the LLM judge. This hybrid approach keeps eval cost proportional to eval complexity.
The regression_source field is populated when a case was added because a production failure was observed. This creates a feedback loop: production failures automatically strengthen the eval harness.
Failure Modes
Eval case overfit to known failures: If eval cases are only added reactively, the suite becomes a regression test for observed failures rather than a behavioral specification. Reserve 30% of the eval development budget for proactive case generation — cases that exercise the intended behavior, not cases that reproduce observed failures.
Judge model prompt instability: If the judge prompt changes between evals, score comparisons between runs are meaningless. The judge prompt is part of the measurement methodology and must be version-controlled with the same rigor as the system prompt.
Coverage collapse: Teams that start with 50 eval cases and never expand them end up with a suite that covers the happy path and nothing else. A production system should have coverage across the behavioral categories the domain expert identified, not just the categories the engineer remembered to test.
Cost-driven avoidance: Running 500 eval cases on every PR at $0.04 each costs $20. Some teams treat this as expensive. The comparison is not $20 vs. free — it is $20 vs. the cost of a production incident, which in my experience runs $10K-$100K in engineering time and customer impact.
Decision Criteria
The minimum viable eval harness for a production LLM system is: 50 cases, a judge rubric, CI integration, and a production sampling pipeline. Below 50 cases you don't have coverage. Above 500 cases you have diminishing returns on new cases versus expanding the judge rubric for existing cases.
The right investment sequence: build the harness before you write the first production prompt, not after the first incident. The cost differential is one sprint of upfront investment versus weeks of reactive debugging.
Closing
The measurement gap in LLM development is the gap between "I think this works" and "I can demonstrate this works at a specified quality threshold across a representative distribution of inputs." Closing that gap is what evals do. The teams that have closed it aren't doing anything architecturally novel — they're applying the same test-before-ship discipline that any mature software team applies, adapted for the probabilistic nature of LLM outputs. The teams that haven't closed it are going to keep describing their systems as "inconsistent" — and they will be right.
Production Readiness Checklist
Before deploying an LLM system with an eval-driven workflow:
• [ ] Minimum 50 eval cases in version control before first production deployment
• [ ] Judge prompt is version-controlled and locked — changes require PR review
• [ ] CI pipeline blocks merge on >5% regression in any eval category
• [ ] Production sampling is running at 2-5% of real requests
• [ ] Failing production samples are automatically added to the regression suite
• [ ] Eval suite covers at least one adversarial case per behavioral category
What I Would Build Differently
The piece I most consistently underestimate is the eval case maintenance burden. When a model is updated, or when the system prompt changes significantly, some fraction of previously-passing eval cases will fail for legitimate reasons — not because the system regressed, but because the intended behavior changed. Teams that don't have a process for distinguishing "this case failed because the system got worse" from "this case failed because the intended behavior changed" end up either ignoring the signal (dangerous) or treating every failing case as a regression (which kills the CI pipeline with false positives). The solution is a quarterly eval review that explicitly re-certifies the ground truth of each case against the current intended behavior. It takes half a day. It is almost never scheduled.

Figure 4. Four-phase pipeline diagram (Behavioral Specification, Prompt Development, CI Integration, Production Monitoring) shown as horizontal workflow with feedback arrows. Below, the eval case JSON schema…
REFERENCES
1. Evals are All You Need — Hamel Husain. hamel.dev (2024).
https://hamel.dev/blog/posts/evals/
2. OpenAI Evals Framework. GitHub / OpenAI (2023).
https://github.com/openai/evals
3. Anthropic Model Card Methodology. Anthropic (2024).
https://www.anthropic.com/model-card
4. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv / NeurIPS (2023).
https://arxiv.org/abs/2306.05685
5. RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv (2023).
https://arxiv.org/abs/2309.15217
6. Holistic Evaluation of Language Models (HELM). Stanford CRFM (2023).
https://crfm.stanford.edu/helm/
7. Promptfoo: Testing and Evaluation for LLM Applications. promptfoo.dev (2024).
https://promptfoo.dev/docs/intro
8. EleutherAI Language Model Evaluation Harness. GitHub / EleutherAI (2023).


