Kimi K2.7 Code

Moonshot

Rank #2 of 11 · 30 PRs · 36/95 golden bugs found

43.1

F1 · +5.6 vs avg

Recall
37.9%+2.7 vs avg
Precision
50.0%+5.5 vs avg
Cost / PR
$0.550
Cost / bug found
$0.46

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):28.1–48.7points. This measures sampling variance from which PRs are in the set.

Judge noise (same submission, 3 independent re-scores): recall37.5 ± 0.6pts (n=3). This is the judge alone — the exact same findings, scored again.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 8Medium 57High 40Critical 6

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug98Performance2Security6Other6

By repository

RepoRecallPrecisionGoldensPRs
cal.com34.8%49.5%236
Discourse50.0%39.0%206
Sentry26.3%54.2%196
Keycloak17.6%31.3%176
Grafana62.5%81.3%166

Per-PR breakdown (30)