Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
NVIDIA announces Nemotron 3 Nano Omni, a new omni-modal understanding model that extends the Nemotron multimodal line to handle text, image, video, and audio inputs. The model achieves top accuracy on several benchmarks including MMlongbench-Doc, OCRBenchV2, WorldSense, DailyOmni, and VoiceBench, and offers up to 9x higher throughput and 2.9x faster reasoning speed compared to alternatives. It uses a hybrid Mamba-Transformer Mixture-of-Experts backbone with C-RADIOv4-H vision and Parakeet-TDT-0.6B-v2 audio encoders, and checkpoints are available in BF16, FP8, and NVFP4 formats on HuggingFace.
From the source
NVIDIA Nemotron 3 Nano Omni is a new omni-modal understanding model built for real-world document analysis, multiple image reasoning, automatic speech recognition, long audio-video understanding, agentic computer use, and general reasoning.
huggingface.co