#llm-evaluation
Release Discipline
One run, many axes — and no winner we can't defend
A Sweep varies model, prompt, language, and tenant over one suite and lays them out in a single matrix. The hard part isn't the cartesian — it's refusing the green 'winner' that's really just noise. This post walks through running a Sweep: pick your axes, read the cost gate, and read a matrix that calls a tie a tie. Built and merged, not yet deployed.
#release-discipline#eval-studio#llm-evaluation#agent-testing#model-comparison
2026-06-18
Release Discipline
Evals for the QA who owns quality but doesn't write YAML
More teams are shipping agentic features than know how to test them — and the one tool that can, promptfoo, makes you write evals as code. Eval Studio brings that capability to the person whose job is quality, not the person who happens to write Python. This is the first post in a series on how we built it.
#release-discipline#eval-studio#llm-evaluation#agent-testing#no-code
2026-06-18