Harnesses find more verified bugs, and open-weight models make them cheap. A systematic evaluation of AI code review: prompts, models, effort levels, and harnesses. Measured, not guessed.
Same pull requests. Same judges. One-shot prompt vs. two agentic review harnesses: compound engineering (parallel persona subagents) and metareview (deterministic gates + adversarial LLM lenses).
…and 1/38 per review, including input, output, and caching costs—even though the harness uses about 1.5× as many tokens.
Two codebases, the hardest slice of a 50-PR benchmark. A deliberate stress test, not a random sample of your repo.
Our intervals cover which PRs we sampled. Run-to-run model variance isn’t in them.
Opus · CE · medium leads at F2′ 0.641; best eligible one-shot is Fable · medium at 0.427. Their paired difference is +0.214 [0.116, 0.275], but selection makes it optimistic.
Frontier model × many-pass harness was beyond budget. Those cells are gaps. Run them, send the data, and we’ll amend the report.