DeepSeek V4 Pro

DeepSeek

Rank #1 of 11 · 30 PRs · 42/95 golden bugs found

43.9

F1 · +6.4 vs avg

Recall
44.2%+9.0 vs avg
Precision
43.6%-0.9 vs avg
Cost / PR
$0.300
Cost / bug found
$0.21

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):32.3–56.3points. This measures sampling variance from which PRs are in the set.

Judge noise (same submission, 3 independent re-scores): recall43.9 ± 0.6pts (n=3). This is the judge alone — the exact same findings, scored again.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 29Medium 50High 22Critical 0

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug97Performance2Security2Other0

By repository

RepoRecallPrecisionGoldensPRs
cal.com52.2%46.1%236
Discourse35.0%31.9%206
Sentry47.4%73.6%196
Keycloak11.8%14.7%176
Grafana75.0%70.8%166

Per-PR breakdown (30)