Muse Spark 1.3

Meta

Rank #10 of 13 · 30 PRs · 37/95 golden bugs found

36.3

F1 · -1.4 vs avg

Recall
39.0%+3.4 vs avg
Precision
34.0%-10.1 vs avg
Cost / PR
$0.530
Cost / bug found
$0.43

Confidence

95% bootstrap interval on recall (2000 resamples over the 30 PRs):28.4–50.6points. This measures sampling variance from which PRs are in the set.

Judge run-to-run variance not yet measured for this model.

Neither measures model run-to-run variance — re-running the review agent itself, not just the judge. One pass per entry.

Run configuration

HarnesskodusAccess pathapiExecution modereplayReasoningvendor-defaultJudgeclaude-haiku-4-5

Severity mix

What the model called its own findings — not recall by severity, goldens aren't severity-tagged.

Low 3Medium 58High 42Critical 3

Category mix

Only bug/performance/security are consistent across models — the rest is free text, bucketed as other.

Bug106Performance0Security0Other0

By repository

RepoRecallPrecisionGoldensPRs
cal.com39.1%37.8%236
Discourse35.0%25.0%206
Sentry31.6%43.1%196
Keycloak23.5%19.0%176
Grafana68.8%62.7%166

Per-PR breakdown (30)