From the source
This blog post investigates unexpectedly low DROP benchmark scores on the Open LLM Leaderboard, identifying issues with normalization and the use of '.' as an end-of-generation token, and proposes using '\n' instead.
From the source
From the source
This blog post investigates unexpectedly low DROP benchmark scores on the Open LLM Leaderboard, identifying issues with normalization and the use of '.' as an end-of-generation token, and proposes using '\n' instead.
From the source
We did a deep dive to understand what was going on, come with us to see what we found out!
huggingface.co