From the source
But which model gives the best results for the cost?
The answer matters: it's the difference between catching a null pointer dereference in production and missing it entirely, and between spending $0.15 per PR and $5.63.
We built a benchmark to find out.
We tested 13 models across 50 real pull requests from five major open-source projects: Sentry, Grafana, Keycloak, Discourse, and Cal.com.
Each model reviewed every PR at least three times, using the same prompts, the same methodology, and the same reasoning effort level ("high" across the board).
A human-curated golden set of known bugs served as ground truth, and an LLM judge scored each review against it.
One question: which model finds the most real bugs per dollar?
Try it yourself: Run /install-code-review in Droid to configure PR reviews in GitHub or GitLab. /review is also available as a standalone skill.
The Money Chart The best model isn't the most expensive one.
Not even close.
F1 Score vs.
Cost per PR Higher is better on Y-axis, lower is better on X-axis.
The best models live in the top-left.
…





