8 models × 3 harnesses × 3 effort levels on the six severity-hardest Martian-benchmark PRs · 66 complete cells · data freeze 2026-09-16 09:35 · quality panels use the true golden set (42 goldens + 105 verified defects with archived reproduction/fix-test evidence = 147 true bugs — REPORT.md §3.1; catalogue: GOLD_DEFECT_CATALOG.md)
Filters — one key area, applies to every panel below
models
harnesses
connections
effort levels
Untick to hide; every chart, ladder, bar, and dumbbell below redraws instantly. Default: all on.
This box sticks to the top as you scroll, shrinks, and slides to the LEFT into the page margin when there is room (it slides back when you scroll up).
Matched framework comparison: across 21 identical model/effort cohorts, mean F2′ is MRV 0.444, CE 0.399, and vanilla 0.265. MRV leads CE in 18/21 point comparisons; paired 95% intervals favor MRV in 8, CE in 3, and include zero in 10. These averages differ from the individual configurations plotted below.
Quality panels score F2′: real bugs found on the true golden set (42 goldens + 105 verified defects), with credit for accepted useful advisories and penalties for unsupported claims, style-only findings, and vague concerns. F2′ = (5T + A) / (4D + T + A + H), where T = distinct true bugs found, D = true-set size, A = accepted useful advisories, and H = effective unsupported findings. All 403 selected reviews have completed adjudication. Advisory credit requires classifier confidence ≥0.70; unsupported-finding penalties require ≥0.80. Findings below the applicable cutoff remain classified but unscored. Confidence is model-reported, not a calibrated probability. Rankings require complete six-PR coverage. The F2′ advisory instrument is recorded separately from legacy bug-matching judges. Verified bug assignments fix bug credit and prevent those findings from receiving advisory credit or unsupported penalties. Every point is one (model × harness × effort level) cell, one selected healthy scored run per PR (n=6). Bars/whiskers are 95% cluster-bootstrap CIs over PRs. Hover for detail; click a point for the full cell card (right); double-click a legend entry to isolate a model; single-click to toggle. Full report: REPORT.md · coverage: COVERAGE.md.
1a · Value for money — how many real bugs does a dollar per PR review buy?
How to read the y-axis — F2′, our overall quality score. It rewards finding real bugs and gives additional credit for accepted useful advisories. The true golden set contains 42 Martian goldens + 105 verified defects (147 total). F2′ = (5T + A) / (4D + T + A + H): T is distinct true bugs found, D is true-set size, A is accepted important non-bug advisories from the merged full review, and H is effective unsupported findings (including false claims, style-only findings, and vague concerns). Advisories add credit without changing bug recall. Legacy F1′ and adjP′ retain their burden-based definitions. All selected reviews are adjudicated. Only useful advisories at confidence ≥0.70 receive credit; unsupported findings at ≥0.80 receive penalties. Lower-confidence findings are classified but unscored. F2′ is a custom score, not measured developer utility. Higher is better; the axis fits the observed scores and confidence intervals. Full definition: REPORT.md §3.1.
x = metered $ per PR review (log) — what a single review costs. One point per complete six-PR cell; all selected reviews are adjudicated. Color = model; shape = harness (○ van, □ CE, △ MRV); connecting lines follow one model × harness across effort levels (solid vanilla, dashed CE, dotted MRV). Hover for the value + CI; click a point for its card below; double-click a legend entry to isolate a model. CIs are off by default.
Takeaway: compare the upper-left cells for strong bug quality plus advisory credit at lower cost. Use the cell cards for observed advisory counts and uncertainty; these panels compare complete six-PR cells.
Cell details
Click a point.
1b · Does a slower setup buy a better review?
A setup that makes you wait longer is only worth it if it reviews better. Every point is one setup; the further left, the less time you spend waiting for the review. When two setups sit at the same height they have the same observed F2′ point estimate; the one further left has lower measured latency. Compare their confidence intervals, then pilot the candidates and measure developer triage time. The target is up and to the left. A point that is high but far right is thorough at the cost of wall-clock rather than dollars — and for a person waiting on a PR, minutes often matter more than cents. y = F2′, our overall quality score — bug quality plus accepted advisory credit, penalised for unsupported findings. Full definition under 1a.
One point per complete six-PR cell; all selected reviews are adjudicated. Color = model; shape = harness (○ van, □ CE, △ MRV); connecting lines follow one model × harness across effort levels (solid vanilla, dashed CE, dotted MRV). Hover for the value + CI; click a point for its card below; double-click a legend entry to isolate a model. CIs are off by default.
Takeaway: GLM Flash and Vision MRV at low effort take about 79 and 95 seconds per review, versus about 100 seconds for Opus MRV low. Several frontier one-shot cells are faster still. GLM medium/high-effort cells can take much longer; compare the individual configurations and their uncertainty. These are measured end-to-end latencies, and this experiment does not isolate the cause of the differences.
Cell details
Click a point.
1c · What does it cost to catch one true bug?
This is the price per true bug — one of the 147 real bugs in our golden set (42 Martian goldens plus 105 verified against archived evidence) — plotted against overall review quality. Cheap and high is the win. A setup far to the left but low is buying cheap catches while missing a lot; one far right and high is thorough but you pay for it. Compare it with 1d, and mind the denominators: 1c counts each distinct true bug once (a bug is credited once no matter how many times it was reported), while 1d divides by findings labelled real by the original campaign model adjudicators — the same bug reported repeatedly counts each time, and beyond-gold findings count too — so the two cost measures use different judgments and denominators. A wide 1c↔1d gap means many campaign model-adjudicated findings per distinct credited bug. y = F2′, our overall quality score — bug quality plus accepted advisory credit, penalised for unsupported findings. Full definition under 1a.
One point per complete six-PR cell; all selected reviews are adjudicated. Color = model; shape = harness (○ van, □ CE, △ MRV); connecting lines follow one model × harness across effort levels (solid vanilla, dashed CE, dotted MRV). Hover for the value + CI; click a point for its card below; double-click a legend entry to isolate a model. CIs are off by default.
Takeaway (true golden set): per distinct true bug found, the GLM harness cells pay far less than the frontier models: glm-flash · MRV · low $0.0020/bug, glm-vis · MRV · low $0.0185/bug, glm-vis · MRV · high $0.204/bug, versus opus · CE · low $0.272/bug and fable · vanilla · high $0.152/bug — and fable finds 51 real bugs to glm-vis MRV high's 76.
Cell details
Click a point.
1d · Cost per campaign model-adjudicated finding
This panel retains the original campaign’s finding-cost accounting.x uses original campaign model judgments and includes repeated reports; y uses revised F2′. The x-axis divides the run's cost by every finding the original campaign model adjudicators labelled real — golden hits plus beyond-gold real findings, counted with multiplicity: if a setup reports the same bug three times, it counts three times. That is a different denominator from panel 1c, which prices each of the 147 distinct true bugs (42 goldens + 105 verified hidden-gold) once. For glm-vis · MRV · low the same run is $0.0037 per campaign model-adjudicated finding (360 reports, including repeats) but $0.019 per distinct true bug (72 credited) — the two numbers are 5× apart and answer different questions: 1d asks 'what does a campaign model-adjudicated finding cost', 1c asks 'what does a distinct bug cost'. Cheap and high is best. A wide 1c↔1d gap means many campaign model-adjudicated findings per distinct bug — including duplicates and beyond-gold items the true-set score does not credit. This legacy denominator is not the revised deduplicated T + A. y = F2′, our overall quality score — bug quality plus accepted advisory credit, penalised for unsupported findings. Full definition under 1a.
One point per complete six-PR cell; all selected reviews are adjudicated. Color = model; shape = harness (○ van, □ CE, △ MRV); connecting lines follow one model × harness across effort levels (solid vanilla, dashed CE, dotted MRV). Hover for the value + CI; click a point for its card below; double-click a legend entry to isolate a model. CIs are off by default.
Takeaway: counting findings labelled real by the original campaign model adjudicators (golden + beyond-gold, with multiplicity), $ per campaign model-adjudicated finding runs $0.0002–$0.37 across cells — the GLM harness cells sit at the cheap end and the high-effort frontier-model cells at the expensive end. Do not read this as the price of a distinct true bug — that is panel 1c, 5× higher for the recommended cell. The frontier panels show which cells offer the strongest observed quality at each price; six-PR scope and classifier uncertainty apply.
Cell details
Click a point.
1e · Do extra tokens buy a better review?
More tokens should mean a better review, or they are just a bigger bill. Every point is one setup; the further left, the fewer tokens a review consumes (cached reads and writes included). If a setup using fewer tokens sits as high as one using more, their observed F2′ point estimates match; this alone does not establish equivalent review quality. Compare confidence intervals and validate developer triage time in a pilot. Up and to the left is best. This is where harness design shows itself: the same model wrapped in a harness usually spends several times the tokens of a single pass, and this panel shows how token use varies alongside observed F2′; a pilot can test whether those differences improve developer outcomes. y = F2′, our overall quality score — bug quality plus accepted advisory credit, penalised for unsupported findings. Full definition under 1a.
One point per complete six-PR cell; all selected reviews are adjudicated. Color = model; shape = harness (○ van, □ CE, △ MRV); connecting lines follow one model × harness across effort levels (solid vanilla, dashed CE, dotted MRV). Hover for the value + CI; click a point for its card below; double-click a legend entry to isolate a model. CIs are off by default.
Takeaway: harness cells burn a median of ~10× the tokens of a vanilla single pass (matched pairs range from ~1× to ~76×); the frontier harnesses are 75–90% cached reads at 10% of list input (why their blended rate drops to 0.14–0.18 ¢ per k-token), while the GLM harnesses use 100–570k tokens/run at list input rates — both arrive cheap per token, ~13× apart in $ per task.
Cell details
Click a point.
1f · How many real bugs does each setup actually find — and how many does the benchmark ignore?
This is the “how much does it actually find” panel — with the benchmark’s blind spot exposed. Each bar is one setup. The grey part is bugs the strict benchmark knows about and scores; the coloured part is real bugs the benchmark never scores — defects verified against archived reproduction and fix-test evidence; some catalogue entries share a test or container-level evidence. A long coloured section means the setup found genuine problems that a benchmark-based score simply cannot see. Read this ranked by bugs found, not by F2′: a setup can lead here and still sit mid-table on quality, because F2′ also rewards accepted useful advisories and penalises unsupported findings. Advisory completeness is tracked separately from the original bug-matching instrument. The two panels answer different questions: 1f asks “how much did it find?”, the quality panels ask “how good is what it found?”
One bar per setup, sorted by total real bugs found. Grey = Martian goldens (the only thing the strict benchmark scores); coloured = hidden-gold defects it never scores. The dashed line at 147 is the observed reference set on the six PRs (42 goldens + 105 verified defects); the small grey tick on a row is that setup’s coverage ceiling — some setups cover fewer than the six PRs and cannot reach 147. Hover for details.
2 · Does more effort buy a better review — and where does it stop paying off?
Every line follows one model and harness as you turn its effort setting up (low → medium → high). If the line climbs, the higher effort has a higher observed F2′; if it flattens or falls, the extra effort has not improved that point estimate. Compare paired uncertainty and pilot developer outcomes before choosing a setting. These six-PR observations do not establish a causal benefit or harm from effort. Compare lines against each other to see whether a cheap model at high effort beats an expensive one at low effort — which is the question most people actually have.
One line per model × harness with points at low → medium → high (left to right). x = metered $ per run (log) — each gridline to the right costs roughly 10× more; y = F2′, our overall quality score (defined under 1a). Colour = model; marker shape = harness (○ vanilla, □ CE, △ MRV). The harness dropdown isolates one harness at a time. The y-axis fits the observed scores and their confidence intervals so the differences remain visible. Hover a point for its value and CI.
3 · Where does a review's token bill actually come from?
Two setups can deliver similar review quality at very different cost, and this panel shows where the money went. Each bar breaks a review's tokens into the four things you pay for: input read for the first time, input re-read from cached context, cache writes paid to store that context, and the model's own output including its reasoning. A setup that is expensive because of output is doing a lot of thinking; one that is expensive because of cache writes is paying to remember things it may never reuse. Read this when you want to change a cost, not just compare it — it tells you which lever to pull.
The report export shows low effort; use this dashboard’s effort dropdown for medium or high. One stacked bar per cell = mean k-tokens per PR across the six PRs, split into fresh input, cached read, cache write and output (incl. reasoning). Label = model · harness · effort level. Default sort: total tokens, descending. Use the effort dropdown to compare like with like, and the sort dropdown to group rows by model or by harness. The global effort filter above also applies here: the panel shows the intersection of the global filter and the local dropdown selection.
4 · How does original-gold scoring change between the top six and the broader run set?
This compares original benchmark goldens for setups with broad coverage. Each dot is a setup’s top-six score minus its score across its available 40–50 PRs. Positive gaps mean the six-PR score is higher. These paired descriptive comparisons do not establish out-of-sample rankings or validate the 147-bug reference set, which was constructed only for the six selected PRs.
One row per cell with at least 40 of 50 PRs covered. Both recall and F1 use the original-gold benchmark lens. x = Δ (top-six − broader run set); the dashed line is zero. Rows are sorted by |Δ|. Use the metric dropdown to switch between recall and F1.
5 · How were the six pull requests selected?
The six were selected for high golden-comment severity weight. Each dot shows one PR’s severity weight and original-gold recall for the selected setup. The highlighted six average severity weight 18.2, compared with 6.1 for the other 44 benchmark PRs; the lowest selected severity ties the highest unselected severity. This describes a deliberately difficult slice. It does not establish how the six-PR true-gold results generalize to other PRs.
x = summed severity weight of original golden comments (Critical = 4, High = 3, Medium = 2, Low = 1). y = per-PR original-gold recall. The dropdown offers setups covering at least 40 PRs; only their available runs are plotted. Highlighted dots = selected PRs; grey = other covered PRs. Dashed lines show unweighted mean per-PR recall within each group, distinct from the pooled recall comparison in panel 4.