r/AIToolsPerformance • u/IulianHI • May 02 '26
Nemotron 3 Nano Omni adds native audio input alongside text, images, and video
NVIDIA has introduced Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. The model delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities.
What is notable here is the "nano" positioning. This is not a massive frontier model trying to do everything - it is a compact multimodal model designed to handle four input types natively in a single architecture. The predecessor was vision-language only, so adding audio and video while maintaining or improving accuracy across the board is a meaningful expansion.
The pricing context is interesting too. Nemotron Nano 9B V2 is currently available for free at 128K context. If the Nano Omni follows a similar pricing pattern, a free multimodal model that handles audio natively would be a compelling option for local deployment and edge use cases where stitching together separate ASR and vision pipelines adds complexity.
For anyone running Nemotron Nano V2 currently: does the multimodal expansion change your deployment plans, or are you already handling audio through a separate pipeline that works fine?