Oossa

NVIDIA releases open Nemotron 3 Nano Omni 30B multimodal model

The 30‑billion‑parameter model can handle text and images now, with video and audio support planned.

By Published by Oossa: 1 min read

National Cancer Institute · Unsplash

NVIDIA has published a new open‑source AI model called Nemotron 3 Nano Omni. It has 30 billion total parameters, but only 3 billion are active at any time thanks to a Mixture‑of‑Experts design. The model can reason over text and images today. For video or audio inputs you must turn off the thinking mode with a special flag.

What the model can do

The model is multimodal, meaning it can process more than one type of data. At launch it supports text‑to‑text and text‑to‑image tasks, such as answering questions about a picture or describing an image. NVIDIA notes that video and audio inputs are technically possible, but you need to set `enable_thinking: false` in the request to avoid errors.

Why it matters

For developers and researchers, Nemotron 3 Nano Omni offers a freely available, large‑scale model that can be fine‑tuned for apps that need both language and visual understanding. Because only a fraction of the 30 B parameters are active, it runs faster and cheaper than a full 30 B model, which could lower the cost of building multimodal products.

Why it matters

Developers can now experiment with a large multimodal model without paying for a commercial API. The active‑expert design keeps compute costs lower, making it feasible for smaller teams to build apps that understand both words and pictures.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.