From the source
This blog post investigates whether foundation models like GPT-4 can be used to evaluate other models' outputs by comparing their labels to human annotations.
The authors curated a test set of prompts and completions from open-source models (Koala 13b, Vicuna 13b, OpenAssistant 12b, Dolly 12b) and collected both human preferences via Scale AI and GPT-4 evaluations.
They expand the Hugging Face Open LLM Leaderboard to include automated benchmarks, human labels, and GPT-4 evaluations.



