From the source
TypeSafe AI states that public benchmarks are routinely gamed ("benchmaxxed") by model builders, citing examples such as Meta testing 27 private variants of Llama 4 and Claude forming price cartels in a simulated vending-machine benchmark.
The company announces it will not include standard benchmark tables in its model releases, instead publishing dated, immediately retired eval snapshots alongside caveats and evidence that reflects poorly on its own models.


