# Hugging Face — Evaluating Audio Reasoning with Big Bench Audio

- Company: Hugging Face (huggingface.co)
- Announced: 2024-12-20T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Models affected: GPT-4o Realtime Preview (Oct '24), GPT-4o Realtime Preview (Dec '24), GPT-4o mini Realtime Preview (Dec '24), GPT-4o ChatCompletions Audio Preview, GPT-4o (Aug '24), Gemini 1.5 Flash (May '24), Gemini 1.5 Flash (Sep '24), Gemini 1.5 Pro (May '24), Gemini 1.5 Pro (Sep '24), Gemini 2.0 Flash (Experimental)
- Source: https://huggingface.co/blog/big-bench-audio-release
- Record: https://forck.live/items/1777-evaluating-audio-reasoning-with-big-bench-audio

Artificial Analysis releases Big Bench Audio, a new evaluation dataset for assessing audio reasoning, and presents benchmark results for GPT-4o and Gemini 1.5 series models across speech and text modalities, revealing a significant speech reasoning gap.

## Evidence

Verbatim from https://huggingface.co/blog/big-bench-audio-release:

> Artificial Analysis is releasing Big Bench Audio, a new evaluation dataset for assessing the reasoning capabilities of audio language models.

---

Record: https://forck.live/items/1777-evaluating-audio-reasoning-with-big-bench-audio
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
