GLM-5.3 Flash
ZhipuRank #5 of 11 · 30 PRs · 37/95 golden bugs found
40.0
F1 · +2.5 vs avg
Recall
39.0%+3.7 vs avg
Precision
41.0%-3.4 vs avg
Cost / PR
$0.080
Cost / bug found
$0.07
Confidence
95% bootstrap interval on recall (2000 resamples over the 30 PRs):29.1–49.4points. This measures sampling variance from which PRs are in the set.
Judge run-to-run variance not yet measured for this model.
Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.
Run configuration
HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5
Severity mix
What the model called its own findings — not recall by severity, goldens aren't severity-tagged.
Low 35Medium 37High 19Critical 4
Category mix
Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.
Bug82Performance6Security7Other0