# NVIDIA Launches Nemotron 3 Nano Omni Model
**作者**: NVIDIA
**日期**: 2026-04-28T16:06:21.000Z
**来源**: [https://x.com/nvidia/status/2049158286461804556](https://x.com/nvidia/status/2049158286461804556)
---

Source: NVIDIA Blog
BYLINE: Kari Briski
Many AI agent systems today juggle separate models for vision, speech and language — losing time and context as they pass data from one model to the other.
NVIDIA today introduced Nemotron 3 Nano Omni, an open multimodal model that combines these capabilities into one, enabling agents to deliver faster, smarter responses with advanced reasoning across video, audio, image and text.
This best-in-class model gives enterprises and developers a production path for more efficient and accurate multimodal AI agents with full deployment flexibility and control.
Nemotron 3 Nano Omni sets a new efficiency frontier for open multimodal models with leading accuracy and low cost, topping six leaderboards for complex document intelligence, and video and audio understanding.
AI-driven companies already adopting Nemotron 3 Nano Omni include Aible, Applied Scientific Intelligence (ASI), @ekacareHQ, @hcompany_ai, and Pyler, with @Amdocs, @Dell, @Docusign, @Infosys, @IQVIA_global , @k_dense_ai, Lila, @Oracle, @PalantirTech, @Quantiphi, @TCS and Zefr evaluating the model.
“To build useful agents, you can’t wait seconds for a model to interpret a screen,” said Gautier Cloix, CEO of H Company. “By building on Nemotron 3 Nano Omni, our agents can rapidly interpret full HD screen recordings — something that wasn’t practical before. This isn’t just a speed boost: It’s a fundamental shift in how our agents perceive and interact with digital environments in real time.”
Nemotron 3 Nano Omni Enables Faster, Leaner Multimodal Agents
Consider an AI agent for customer support processing a screen recording while analyzing uploaded call audio and checking data logs — or an agent for finance tasked with parsing PDFs, spreadsheets, charts and voice notes. Today, most agentic systems accomplish these tasks with separate models for vision, speech and language.
This approach increases latency through repeated inference passes, fragments context across modalities, and adds cost and inaccuracies over time.
By combining vision and audio encoders within its 30B-A3B, hybrid mixture-of-experts architecture, Nemotron 3 Nano Omni eliminates the need for separate perception models, driving inference efficiency at scale. Using this model, an AI system will achieve 9x higher throughput than other open omni models with the same interactivity, resulting in lower cost and better scalability without sacrificing responsiveness.
In agentic systems, Nemotron 3 Nano Omni can work alongside proprietary cloud models or other NVIDIA Nemotron open models — such as Nemotron 3 Super for high-frequency execution or Nemotron 3 Ultra for complex planning — as well as proprietary models from other providers, to power sub-agents for agentic workflows such as computer use, document intelligence and audio-video reasoning.
- Computer use agents — Nemotron 3 Nano Omni powers the perception loop for agents navigating graphical user interfaces, reasoning over onscreen content and understanding user interface state over time. H Company’s latest computer usage agent, powered by Nemotron 3 Nano Omni, uses a native input resolution of 1920x1080 pixels to achieve high-fidelity visual reasoning. In preliminary evaluations on the OSWorld benchmark, this integration showed a significant leap in navigating complex graphical interfaces and used Nemotron 3 Nano Omni’s ability to process very high-resolution images.
- Document intelligence — Interprets documents, charts, tables, screenshots and mixed-media inputs, enabling agents to reason across visual structure and text content coherently. Critical for enterprise analysis and compliance workflows.
- Audio and video understanding — For customer service, research and monitoring workflows, Nemotron 3 Nano Omni maintains audio-video context, tying what was said, shown and documented into a single reasoning stream instead of disconnected summaries

Open and Customizable, Deployable Anywhere
Nemotron 3 Nano Omni is released with open weights, datasets and training techniques — giving organizations full transparency and control over how the model is customized and deployed.
Developers can use tools like NVIDIA NeMo for customization, evaluation and optimization for domain-specific use cases. Because the Nemotron family of models is open, organizations can deploy them in environments that meet regulatory, sovereignty or data localization requirements. The Nemotron 3 family — including Nano, Super and Ultra models — has seen over 50 million downloads in the past year. Omni extends the family’s capabilities into multimodal and agentic domains.
The model is available on Hugging Face, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice and through a broad ecosystem of NVIDIA Cloud Partners, inference platforms and cloud service providers.
Its open, lightweight architecture supports consistent deployment from local systems like NVIDIA DGX Spark and DGX Station to data center and cloud environments.
Visit NVIDIA’s technical blog for tutorials, cookbooks and deployment guides for Nemotron 3 Nano Omni use cases.
Stay up to date on agentic AI, NVIDIA Nemotron and more by subscribing to NVIDIA news, joining the community and following NVIDIA AI on LinkedIn, Instagram, X and Facebook.
Explore self-paced video tutorials and livestreams.
At a Glance What it is - An open, omni-modal reasoning model — the highest-efficiency open multimodal model of its kind with leading accuracy
What it handles - Text, images, audio, video, documents, charts and graphical interfaces (input); text (output)
Who it's for - Enterprises and developers building fast and reliable, agentic systems that need a multimodal perception sub-agent
How it works - Functions as the "eyes and ears" in a system of agents, working alongside models like Nemotron 3 Super and Ultra or other proprietary models
Why it matters - Leading multimodal accuracy and 9x higher throughput than other open omni models with the same interactivity, resulting in lower cost and better scalability without sacrificing responsiveness.
Architecture - 30B-A3B hybrid MoE with Conv3D, EVS, 256K context
Availability - April 28th, 2026 via Hugging Face, NVIDIA NIM and 25+ partner platforms
## 相关链接
- [NVIDIA](https://x.com/nvidia)
- [@nvidia](https://x.com/nvidia)
- [74K](https://x.com/nvidia/status/2049158286461804556/analytics)
- [NVIDIA Blog](https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/)
- [topping six leaderboards](https://developer.nvidia.com/blog/nvidia-nemotron-3-nano-omni-powers-multimodal-agent-reasoning-in-a-single-efficient-open-model)
- [Aible](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.aible.com%2Fnemotron3nano-omni-aiagent&data=05%7C02%7Cnhereth%40nvidia.com%7Ca3fd5bfc07a94dc29e9a08dea4090952%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639128555922088829%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=2GVFO7lcp6Vbha2vR%2B7EFJURNiwXcJcfc9Sqn5MgzgM%3D&reserved=0)
- [Applied Scientific Intelligence (ASI)](https://appliedscientific.ai/research/scientific-ai-literature-agent-nvidia-nemotron-nano-omni?utm_source=nvidia-blog)
- [@ekacareHQ](https://x.com/@ekacareHQ)
- [@hcompany_ai](https://x.com/@hcompany_ai)
- [Pyler](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fpyler.tech%2Farticles%2Fscaling-trustworthy-video-safety-with-nvidia-nemotron-3-nano-omni&data=05%7C02%7Cnhereth%40nvidia.com%7Ccfbcee7b93c24b57a80608dea4b3b5ec%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639129288967303883%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=9RgDLwIiHagnD8Sm8vxYqe4kvDsTfK4bQTw%2BtFpAhcY%3D&reserved=0)
- [@Amdocs](https://x.com/@Amdocs)
- [@Dell](https://x.com/@Dell)
- [@Docusign](https://x.com/@Docusign)
- [@Infosys](https://x.com/@Infosys)
- [@IQVIA_global](https://x.com/@IQVIA_global)
- [@k_dense_ai](https://x.com/@k_dense_ai)
- [@Oracle](https://x.com/@Oracle)
- [@PalantirTech](https://x.com/@PalantirTech)
- [@Quantiphi](https://x.com/@Quantiphi)
- [@TCS](https://x.com/@TCS)
- [mixture-of-experts](https://www.nvidia.com/en-us/glossary/mixture-of-experts/)
- [AI system will achieve 9x higher throughput](https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-inteligence)
- [computer usage agent](https://www.youtube.com/watch?v=kSi9JS2l0Ww)
- [NVIDIA NeMo](https://www.nvidia.com/en-us/ai-data-science/products/nemo/)
- [Hugging Face](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16)
- [OpenRouter](https://openrouter.ai/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free)
- [build.nvidia.com](http://build.nvidia.com/)
- [NVIDIA Cloud Partners](https://www.nvidia.com/en-us/data-center/gpu-cloud-computing/partners/)
- [NVIDIA DGX Spark](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)
- [DGX Station](https://www.nvidia.com/en-us/products/workstations/dgx-station/)
- [tutorials, cookbooks and deployment guides](https://developer.nvidia.com/blog/nvidia-nemotron-3-nano-omni-powers-multimodal-agent-reasoning-in-a-single-efficient-open-model)
- [NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/)
- [NVIDIA news](https://www.nvidia.com/en-us/executive-insights/generative-ai-tools/?modal=stay-inf)
- [joining the community](https://developer.nvidia.com/community)
- [LinkedIn](https://www.linkedin.com/showcase/nvidia-ai/posts/?feedView=all)
- [Instagram](https://www.instagram.com/nvidiaai/?hl=en)
- [Facebook](https://www.facebook.com/NVIDIAAI)
- [self-paced video tutorials and livestreams](https://youtube.com/playlist?list=PL5B692fm6--vdRKB14FImVi7MTJ77zjn4&feature=shared)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [12:06 AM · Apr 29, 2026](https://x.com/nvidia/status/2049158286461804556)
- [74.4K Views](https://x.com/nvidia/status/2049158286461804556/analytics)
- [View quotes](https://x.com/nvidia/status/2049158286461804556/quotes)
---
*导出时间: 2026/4/29 08:50:57*
---
## 中文翻译
# NVIDIA 发布 Nemotron 3 Nano Omni 模型
**作者**: NVIDIA
**日期**: 2026-04-28T16:06:21.000Z
**来源**: [https://x.com/nvidia/status/2049158286461804556](https://x.com/nvidia/status/2049158286461804556)
---

来源:NVIDIA 博客
作者:Kari Briski
当今许多 AI 智能体系统都在视觉、语音和语言方面分别使用不同的模型——在模型之间传递数据时会浪费时间并丢失上下文。
NVIDIA 今日发布了 Nemotron 3 Nano Omni,这是一个开放的多模态模型,将这些能力集于一体,使智能体能够针对视频、音频、图像和文本进行高级推理,从而做出更快、更智能的响应。
这款同类最佳模型为企业及开发者提供了一条生产路径,用于构建更高效、准确的多模态 AI 智能体,并具备完整的部署灵活性和控制权。
Nemotron 3 Nano Omni 为开放多模态模型树立了全新的效率标杆,在拥有领先的准确性和低成本的同时,在复杂文档智能、视频和音频理解领域的六个排行榜上名列前茅。
已有 Aible、Applied Scientific Intelligence (ASI)、@ekacareHQ、@hcompany_ai 和 Pyler 等 AI 驱动型企业采用了 Nemotron 3 Nano Omni,同时 @Amdocs、@Dell、@Docusign、@Infosys、@IQVIA_global、@k_dense_ai、Lila、@Oracle、@PalantirTech、@Quantiphi、@TCS 和 Zefr 正在评估该模型。
“要构建有用的智能体,你不能花好几秒钟等待模型去解读屏幕,”H 公司首席执行官 Gautier Cloix 表示,“通过基于 Nemotron 3 Nano Omni 构建,我们的智能体能够快速解读全高清屏幕录制内容——这在以前是不切实际的。这不仅仅是速度的提升:这是我们的智能体实时感知和与数字环境互动方式的根本性转变。”
Nemotron 3 Nano Omni 赋能更快速、更精简的多模态智能体
试想一个用于客户支持的 AI 智能体在处理屏幕录制的同时分析上传的通话音频并检查数据日志——或者一个用于金融的智能体负责解析 PDF、电子表格、图表和语音备忘录。目前,大多数智能体系统都通过针对视觉、语音和语言的独立模型来完成这些任务。
这种方法通过反复的推理传递增加了延迟,导致上下文在不同模态间碎片化,并随时间推移增加成本和不准确性。
通过在其 30B-A3B 混合专家混合架构中结合视觉和音频编码器,Nemotron 3 Nano Omni 消除了对独立感知模型的需求,从而推动了大规模推理效率。使用该模型,AI 系统在保持相同交互性的情况下,其吞吐量将比其他开放全能模型高出 9 倍,从而在不牺牲响应速度的情况下降低成本并提高可扩展性。
在智能体系统中,Nemotron 3 Nano Omni 可以与专有云模型或其他 NVIDIA Nemotron 开放模型(例如用于高频执行的 Nemotron 3 Super 或用于复杂规划的 Nemotron 3 Ultra)以及来自其他提供商的专有模型配合工作,为计算机使用、文档智能和音视频推理等智能体工作流程提供支持。
- 计算机使用智能体 —— Nemotron 3 Nano Omni 为导航图形用户界面的智能体提供感知循环,能够对屏幕内容进行推理并随时间理解用户界面状态。H 公司最新的计算机使用智能体由 Nemotron 3 Nano Omni 驱动,使用 1920x1080 像素的原始输入分辨率来实现高保真视觉推理。在 OSWorld 基准测试的初步评估中,这种集成在导航复杂图形界面方面显示出显著飞跃,并利用了 Nemotron 3 Nano Omni 处理极高分辨率图像的能力。
- 文档智能 —— 能够解读文档、图表、表格、屏幕截图和混合媒体输入,使智能体能够连贯地对视觉结构和文本内容进行推理。对于企业分析和合规工作流程至关重要。
- 音视频理解 —— 对于客户服务、研究和监控工作流程,Nemotron 3 Nano Omni 能够保持音视频上下文,将所说的、展示的和记录的内容关联到单一推理流中,而不是生成断开连接的摘要。

开放、可定制,随处部署
Nemotron 3 Nano Omni 随开放权重、数据集和训练技术一起发布——让组织对模型的定制和部署方式拥有完全的透明度和控制权。
开发者可以使用 NVIDIA NeMo 等工具进行定制、评估和优化,以适应特定领域的用例。由于 Nemotron 系列模型是开放的,组织可以在满足监管、主权或数据本地化要求的环境中部署它们。Nemotron 3 系列——包括 Nano、Super 和 Ultra 模型——在过去一年中已获得超过 5000 万次下载。Omni 将该系列的能力扩展到了多模态和智能体领域。
该模型现已在 Hugging Face、OpenRouter 和 build.nvidia.com 上作为 NVIDIA NIM 微服务提供,并通过广泛的 NVIDIA 云合作伙伴生态系统、推理平台和云服务提供商提供。
其开放、轻量级的架构支持从 NVIDIA DGX Spark 和 DGX Station 等本地系统到数据中心和云环境的一致部署。
访问 NVIDIA 技术博客,获取有关 Nemotron 3 Nano Omni 用例的教程、食谱和部署指南。
通过订阅 NVIDIA 新闻、加入社区并在 LinkedIn、Instagram、X 和 Facebook 上关注 NVIDIA AI,及时了解智能体 AI、NVIDIA Nemotron 等更多信息。
探索自主-paced 视频教程和直播。
**一目了然**
**它是什么** —— 一个开放的全能推理模型——同类中效率最高的开放多模态模型,具有领先的准确性。
**它能处理什么** —— 文本、图像、音频、视频、文档、图表和图形界面(输入);文本(输出)。
**适用对象** —— 构建快速、可靠的智能体系统的企业和开发者,这些系统需要多模态感知子智能体。
**如何运作** —— 充当智能体系统中的“眼睛和耳朵”,与 Nemotron 3 Super 和 Ultra 等模型或其他专有模型配合工作。
**为何重要** —— 领先的多模态准确性,在保持相同交互性的情况下,吞吐量比其他开放全能模型高 9 倍,从而在不牺牲响应速度的情况下降低成本并提高可扩展性。
**架构** —— 30B-A3B 混合 MoE,采用 Conv3D、EVS,256K 上下文。
**可用性** —— 2026 年 4 月 28 日通过 Hugging Face、NVIDIA NIM 和 25+ 合作伙伴平台推出。
## 相关链接
- [NVIDIA](https://x.com/nvidia)
- [@nvidia](https://x.com/nvidia)
- [74K](https://x.com/nvidia/status/2049158286461804556/analytics)
- [NVIDIA Blog](https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/)
- [topping six leaderboards](https://developer.nvidia.com/blog/nvidia-nemotron-3-nano-omni-powers-multimodal-agent-reasoning-in-a-single-efficient-open-model)
- [Aible](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.aible.com%2Fnemotron3nano-omni-aiagent&data=05%7C02%7Cnhereth%40nvidia.com%7Ca3fd5bfc07a94dc29e9a08dea4090952%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639128555922088829%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=2GVFO7lcp6Vbha2vR%2B7EFJURNiwXcJcfc9Sqn5MgzgM%3D&reserved=0)
- [Applied Scientific Intelligence (ASI)](https://appliedscientific.ai/research/scientific-ai-literature-agent-nvidia-nemotron-nano-omni?utm_source=nvidia-blog)
- [@ekacareHQ](https://x.com/@ekacareHQ)
- [@hcompany_ai](https://x.com/@hcompany_ai)
- [Pyler](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fpyler.tech%2Farticles%2Fscaling-trustworthy-video-safety-with-nvidia-nemotron-3-nano-omni&data=05%7C02%7Cnhereth%40nvidia.com%7Ccfbcee7b93c24b57a80608dea4b3b5ec%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639129288967303883%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=9RgDLwIiHagnD8Sm8vxYqe4kvDsTfK4bQTw%2BtFpAhcY%3D&reserved=0)
- [@Amdocs](https://x.com/@Amdocs)
- [@Dell](https://x.com/@Dell)
- [@Docusign](https://x.com/@Docusign)
- [@Infosys](https://x.com/@Infosys)
- [@IQVIA_global](https://x.com/@IQVIA_global)
- [@k_dense_ai](https://x.com/@k_dense_ai)
- [@Oracle](https://x.com/@Oracle)
- [@PalantirTech](https://x.com/@PalantirTech)
- [@Quantiphi](https://x.com/@Quantiphi)
- [@TCS](https://x.com/@TCS)
- [mixture-of-experts](https://www.nvidia.com/en-us/glossary/mixture-of-experts/)
- [AI system will achieve 9x higher throughput](https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-inteligence)
- [computer usage agent](https://www.youtube.com/watch?v=kSi9JS2l0Ww)
- [NVIDIA NeMo](https://www.nvidia.com/en-us/ai-data-science/products/nemo/)
- [Hugging Face](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16)
- [OpenRouter](https://openrouter.ai/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free)
- [build.nvidia.com](http://build.nvidia.com/)
- [NVIDIA Cloud Partners](https://www.nvidia.com/en-us/data-center/gpu-cloud-computing/partners/)
- [NVIDIA DGX Spark](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)
- [DGX Station](https://www.nvidia.com/en-us/products/workstations/dgx-station/)
- [tutorials, cookbooks and deployment guides](https://developer.nvidia.com/blog/nvidia-nemotron-3-nano-omni-powers-multimodal-agent-reasoning-in-a-single-efficient-open-model)
- [NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/)
- [NVIDIA news](https://www.nvidia.com/en-us/executive-insights/generative-ai-tools/?modal=stay-inf)
- [joining the community](https://developer.nvidia.com/community)
- [LinkedIn](https://www.linkedin.com/showcase/nvidia-ai/posts/?feedView=all)
- [Instagram](https://www.instagram.com/nvidiaai/?hl=en)
- [Facebook](https://www.facebook.com/NVIDIAAI)
- [self-paced video tutorials and livestreams](https://youtube.com/playlist?list=PL5B692fm6--vdRKB14FImVi7MTJ77zjn4&feature=shared)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [12:06 AM · Apr 29, 2026](https://x.com/nvidia/status/2049158286461804556)
- [74.4K Views](https://x.com/nvidia/status/2049158286461804556/analytics)
- [View quotes](https://x.com/nvidia/status/2049158286461804556/quotes)
---
*导出时间: 2026/4/29 08:50:57*