Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Gemma 3n, a natively multimodal on-device model supporting image, text, audio, and video inputs, is now fully available in the open-source ecosystem. Two sizes (E2B and E4B) are released with base and instruct variants, achieving competitive performance (e.g., E4B scores 1300+ on LMArena) and efficient memory usage (E2B runs in ~2GB GPU RAM, E4B in ~3GB). The models are supported by transformers, timm, MLX, llama.cpp, and other libraries.
From the source
Gemma 3n was announced as a preview during Google I/O. The on-device community got really excited, because this is a model designed from the ground up to run locally on your hardware. On top of that, it’s natively multimodal, supporting image, text, audio, and video inputs 🤯
huggingface.co