# Hugging Face — Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?

- Company: Hugging Face (huggingface.co)
- Announced: 2024-03-05T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/leaderboard-contextual
- Record: https://forck.live/items/1929-introducing-contextual-how-well-can-your-multimodal-model-jointly-reason-over

Researchers from UCLA introduce ConTextual, a dataset and leaderboard for evaluating multimodal models on context-sensitive text-rich visual reasoning tasks. Initial experiments show that both proprietary and open-source models struggle on this benchmark compared to humans.

## Evidence

Verbatim from https://huggingface.co/blog/leaderboard-contextual:

> That’s why we (researchers from University of California Los Angeles) created ConTextual, a Context-sensitive Text-rich visuaL reasoning dataset for evaluating LMMs. We also released a leaderboard, so that the community can see for themselves which models are the best at this task.

---

Record: https://forck.live/items/1929-introducing-contextual-how-well-can-your-multimodal-model-jointly-reason-over
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
