Qwen3.8 27B

Alibaba

Rank #9 of 11 · 29 PRs · 32/93 golden bugs found

36.2

F1 · -1.3 vs avg

Recall
34.4%-0.8 vs avg
Precision
38.1%-6.3 vs avg
Cost / PR
$0.340
Cost / bug found
$0.31

Confidence

95% bootstrap interval on recall (2000 resamples over the 29 PRs):24.1–45.4points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 27Medium 36High 23Critical 5

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug86Performance3Security0Other8

By repository

RepoRecallPrecisionGoldensPRs
cal.com47.8%32.6%236
Sentry26.3%60.7%196
Discourse27.8%28.0%185
Keycloak17.6%15.3%176
Grafana50.0%67.5%166

Per-PR breakdown (29)