# Hugging Face — Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

- Company: Hugging Face (huggingface.co)
- Announced: 2026-04-28T15:58:57+00:00
- Category: new-model
- Subject: Platform
- Open weights: yes
- Models affected: Nemotron 3 Nano Omni
- Source: https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence
- Record: https://forck.live/items/1510-introducing-nvidia-nemotron-3-nano-omni-long-context-multimodal-intelligence

NVIDIA announces Nemotron 3 Nano Omni, a new omni-modal understanding model that extends the Nemotron multimodal line to handle text, image, video, and audio inputs. The model achieves top accuracy on several benchmarks including MMlongbench-Doc, OCRBenchV2, WorldSense, DailyOmni, and VoiceBench, and offers up to 9x higher throughput and 2.9x faster reasoning speed compared to alternatives. It uses a hybrid Mamba-Transformer Mixture-of-Experts backbone with C-RADIOv4-H vision and Parakeet-TDT-0.6B-v2 audio encoders, and checkpoints are available in BF16, FP8, and NVFP4 formats on HuggingFace.

## Evidence

Verbatim from https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence:

> NVIDIA Nemotron 3 Nano Omni is a new omni-modal understanding model built for real-world document analysis, multiple image reasoning, automatic speech recognition, long audio-video understanding, agentic computer use, and general reasoning.

---

Record: https://forck.live/items/1510-introducing-nvidia-nemotron-3-nano-omni-long-context-multimodal-intelligence
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
