# Hugging Face — Can foundation models label data like humans?

- Company: Hugging Face (huggingface.co)
- Announced: 2023-06-12
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/open-llm-leaderboard-rlhf
- Record: https://forck.live/items/2044-can-foundation-models-label-data-like-humans
- Subject: Platform
- Open weights: yes
- Models affected: Koala 13b, Vicuna 13b, OpenAssistant 12b, Dolly 12b

This blog post investigates whether foundation models like GPT-4 can be used to evaluate other models' outputs by comparing their labels to human annotations. The authors curated a test set of prompts and completions from open-source models (Koala 13b, Vicuna 13b, OpenAssistant 12b, Dolly 12b) and collected both human preferences via Scale AI and GPT-4 evaluations. They expand the Hugging Face Open LLM Leaderboard to include automated benchmarks, human labels, and GPT-4 evaluations.

## Evidence

Verbatim from https://huggingface.co/blog/open-llm-leaderboard-rlhf:

> In this blog post, we'll zoom in on where you can and cannot trust the data labels you get from the LLM of your choice by expanding the Open LLM Leaderboard evaluation suite.

---

Record: https://forck.live/items/2044-can-foundation-models-label-data-like-humans
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
