# Factory AI — Which Model Reviews Code Best?

- Company: Factory AI (factory.com)
- Announced: 2026-04-29
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://factory.com/news/code-review-benchmark
- Record: https://forck.live/items/17823-which-model-reviews-code-best
- Subject: Droid

Every code review Droid produces is backed by a model. But which model gives the best results for the cost? The answer matters: it's the difference between catching a null pointer dereference in production and missing it entirely, and between spending $0.15 per PR and $5.63. We built a benchmark to find out. We tested 13 models across 50 real pull requests from five major open-source projects: Sentry, Grafana, Keycloak, Discourse, and Cal.com. Each model reviewed every PR at least three times, using the same prompts, the same methodology, and the same reasoning effort level ("high" across the board). A human-curated golden set of known bugs served as ground truth, and an LLM judge scored each review against it. One question: which model finds the most real bugs per dollar? Try it yourself: Run /install-code-review in Droid to configure PR reviews in GitHub or GitLab. /review is also available as a standalone skill. The Money Chart The best model isn't the most expensive one. Not even close. F1 Score vs. Cost per PR Higher is better on Y-axis, lower is better on X-axis. The best models live in the top-left. …

---

Record: https://forck.live/items/17823-which-model-reviews-code-best
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
