
Nvidia just dropped Nemotron 3 Nano Omni, an open-source multimodal model that fuses vision, audio, and text into a single model — no more running separate models for each modality in your agent stack. It uses a 30B-A3B mixture-of-experts (MoE) architecture, with open weights, datasets, and training recipes.
The model tops six leaderboards for document intelligence, video, and audio understanding, including MMlongbench-Doc, OCRBenchV2, WorldSense, DailyOmni, and VoiceBench. And here's the kicker: it achieves roughly 9.2x the system throughput of comparable open-source omni models in video reasoning while maintaining the same user interaction latency. For multi-document reasoning, it's about 7.4x faster.
Nemotron 3 Nano Omni supports FP8 and NVFP4 quantization, meaning it can run on everything from Jetson edge devices to full datacenter hardware. Adoption is already underway: H Company, Foxconn, and Palantir are among the early evaluators. The success of Palantir has inspired a wave of imitators in China, including Zhongshu Ruizhi, which recently landed a nine-figure Series B. The Nemotron 3 family has racked up over 50 million downloads in the past year.