NVIDIA announces Llama Nemotron Nano VL, an 8B vision language model optimized for intelligent document processing with high accuracy OCR and multimodal understanding.
Qwen-TTS now supports generating speech in three Chinese dialects: Pekingese, Shanghainese, and Sichuanese, and offers seven Chinese-English bilingual voices.
SGLang now supports Hugging Face Transformers as a backend, allowing any Transformers-compatible model to run with high-performance inference out of the box, with automatic fallback when a model is…
Alibaba introduces Qwen VLo, a unified multimodal understanding and generation model that can both understand images and generate high-quality recreations.