Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face announces FilBench, a comprehensive evaluation suite for assessing LLM performance on Tagalog, Filipino, and Cebuano across four categories: Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. The researchers evaluated more than 20 state-of-the-art models, finding that region-specific LLMs like SEA-LION and SeaLLM are parameter-efficient but still lag behind GPT-4o, and that translation tasks remain challenging for current models.
From the source
Thatβs why we developed FilBench: a comprehensive evaluation suite to assess the capabilities of LLMs for Tagalog, Filipino (the standardized form of Tagalog), and Cebuano, on fluency, linguistic and translation abilities, as well as specific cultural knowledge.
huggingface.co