# OpenAI — MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

- Company: OpenAI (openai.com)
- Announced: 2024-10-10T10:00:00+00:00
- Category: research-paper
- Subject: GPT / ChatGPT / API
- Source: https://openai.com/index/mle-bench
- Record: https://forck.live/items/681-mle-bench-evaluating-machine-learning-agents-on-machine-learning-engineering

Introduction of MLE-bench, a benchmark for evaluating machine learning agents on machine learning engineering tasks.

## Evidence

Verbatim from https://openai.com/index/mle-bench:

> We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.

---

Record: https://forck.live/items/681-mle-bench-evaluating-machine-learning-agents-on-machine-learning-engineering
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
