We chat with GDM's Head of Developer Experience to catch up on Gemma 4, their ultra efficient open source model featuring a new transformer architecture featuring per-layer embeddings, enabling effective parameter offloading where only a fraction (e.g., 2B of 5B parameters) needs to be loaded into the GPU for fast inference, ideal for on-device use. We also chat Gemini Nano, Native Multimodality, and Finetuning and growth trends seen at AIE Europe.