NVIDIA has published a new open‑source AI model called Nemotron 3 Nano Omni. It has 30 billion total parameters, but only 3 billion are active at any time thanks to a Mixture‑of‑Experts design. The model can reason over text and images today. For video or audio inputs you must turn off the thinking mode with a special flag.
What the model can do
The model is multimodal, meaning it can process more than one type of data. At launch it supports text‑to‑text and text‑to‑image tasks, such as answering questions about a picture or describing an image. NVIDIA notes that video and audio inputs are technically possible, but you need to set `enable_thinking: false` in the request to avoid errors.
Why it matters
For developers and researchers, Nemotron 3 Nano Omni offers a freely available, large‑scale model that can be fine‑tuned for apps that need both language and visual understanding. Because only a fraction of the 30 B parameters are active, it runs faster and cheaper than a full 30 B model, which could lower the cost of building multimodal products.
Why it matters
Developers can now experiment with a large multimodal model without paying for a commercial API. The active‑expert design keeps compute costs lower, making it feasible for smaller teams to build apps that understand both words and pictures.