Can foundation models label data like humans?
This blog post investigates whether foundation models like GPT-4 can be used to evaluate other models' outputs by comparing their labels to human annotations. The authors curated a test set of prompts and completions from open-source models (Koala 13b, Vicuna 13b, OpenAssistant 12b, Dolly 12b) and collected both human preferences via Scale AI and GPT-4 evaluations. They expand the Hugging Face Open LLM Leaderboard to include automated benchmarks, human labels, and GPT-4 evaluations.
