# Hugging Face — Improving Prompt Consistency with Structured Generations

- Company: Hugging Face (huggingface.co)
- Announced: 2024-04-30T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/evaluation-structured-outputs
- Record: https://forck.live/items/1899-improving-prompt-consistency-with-structured-generations

Hugging Face's research team, with collaborators at Dottxt, investigates how LLM evaluation results are highly sensitive to prompt format changes, showing that performance can vary by 10 points and model rankings can be inconsistent across different prompt formats. They describe experiments on MMLU subsets with 8 prompt formats and 5 models, highlighting the need for greater prompt consistency.

## Evidence

Verbatim from https://huggingface.co/blog/evaluation-structured-outputs:

> the Leaderboards and Evals research team at Hugging Face did small experiments, which highlighted how fickle evaluation can be. For a given task, results are extremely sensitive to minuscule changes in prompt format!

---

Record: https://forck.live/items/1899-improving-prompt-consistency-with-structured-generations
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
