Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Introduces FutureBench, a benchmark for evaluating AI agents on their ability to predict future events, aiming to measure reasoning and synthesis skills while avoiding data contamination by using events that have not yet occurred.
From the source
We therefore propose evaluating agents on their ability to predict future events (Ye et al., 2024; Karger et al., 2025).
together.ai