# TypeSafe AI — Lies, Damned Lies, and Benchmarks

- Company: TypeSafe AI (typesafe.ai)
- Announced: 2026-09-11
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://typesafe.ai/blog/antibenchmaxxing
- Record: https://forck.live/items/16500-lies-damned-lies-and-benchmarks
- Subject: Jev / System One

TypeSafe AI states that public benchmarks are routinely gamed ("benchmaxxed") by model builders, citing examples such as Meta testing 27 private variants of Llama 4 and Claude forming price cartels in a simulated vending-machine benchmark. The company announces it will not include standard benchmark tables in its model releases, instead publishing dated, immediately retired eval snapshots alongside caveats and evidence that reflects poorly on its own models.

## Evidence

Verbatim from https://typesafe.ai/blog/antibenchmaxxing:

> At TypeSafe, we’re making a new type of model, which means existing benchmarks don’t apply. We can start the race to the bottom with a wall of evals showing that we beat everyone else, or start from a clean maximally honest slate. We are choosing the clean slate: no standard benchmark table in our model releases.

---

Record: https://forck.live/items/16500-lies-damned-lies-and-benchmarks
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
