This blog post describes an experiment testing how well LLMs can fix their mistakes when given feedback in plain English, using a simple calendar API scenario.
Welcome PaliGemma 2 – New vision language models by Google
Hugging Face announces the release of Google's PaliGemma 2 vision language models, which upgrade the text decoder to Gemma 2 and come in 3B, 10B, and 28B parameter sizes with multiple input…
Hugging Face and collaborators announce the AraGen benchmark and leaderboard for Arabic LLMs, introducing the 3C3H evaluation measure that assesses model responses across six dimensions including…
A case study from Capital Fund Management (CFM) leveraging open-source LLMs and the Hugging Face Ecosystem for Named Entity Recognition (NER) on financial data, achieving up to 6.4% accuracy…
Hugging Face published a guide for open source developers about the EU AI Act, explaining its impact on AI systems and models, risk categories, and steps for compliance, including using Hugging Face…
OpenAI published the o1 System Card, detailing safety work, external red teaming and frontier risk evaluations prior to releasing the o1 and o1-mini models.
Sakana AI presents CycleQD, a framework for evolving a population of small LLM agents (8B parameters) using model merging and quality diversity to excel in agentic tasks, supported by the Japanese…