48 / 108 · 03 Agentic Testing Architectures · The Critic-Actor Pattern← prev⊞ allnext →☰ Read as one page
6.7Quality Benchmarks
In benchmarks, critic-actor pipelines produce measurably better tests:
| Metric | Single-Agent Generation | Critic-Actor Pipeline | Improvement |
|---|---|---|---|
| Tautology rate | 15-20% | 2-5% | 75% reduction |
| Mutation score | 55-65% | 70-80% | +15-25 points |
| Coverage gaps identified | 0 (by definition) | 3-5 per spec | N/A |
| Assertion density (per test) | 1.5 | 3.2 | 2x |
| Review time (human) | 30 min / 30 tests | 10 min / 30 tests | 3x faster |
The human review time drops because the Critic has already caught the most common issues (tautologies, weak assertions, missing scenarios).