# Hugging Face — ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

- Company: Hugging Face (huggingface.co)
- Announced: 2026-06-30T18:32:50+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/ibm-research/scarfbench
- Record: https://forck.live/items/1468-scarfbench-benchmarking-ai-agents-for-enterprise-java-framework-migration

IBM Research introduces ScarfBench, an open benchmark for evaluating AI agents on enterprise Java framework migration tasks. The benchmark includes 34 applications, 102 framework implementations, 204 migration tasks, about 151K lines of code, and 1,331 expert-written tests. Evaluations show that even the strongest agents achieve less than 10% behavioral success, agents are overconfident in reporting build success (Claude Code reported 29/30 builds successful but only 22 actually built), migration is iterative, configuration dominates effort, and environment and tooling matter. The leaderboard is available at scarfbench.info/leaderboard.

## Evidence

Verbatim from https://huggingface.co/blog/ibm-research/scarfbench:

> ScarfBench (Self-Contained Application Refactoring Benchmark), an open benchmark for evaluating AI agents on cross-framework migration tasks in Enterprise Java.

---

Record: https://forck.live/items/1468-scarfbench-benchmarking-ai-agents-for-enterprise-java-framework-migration
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
