NVIDIA Developer Blog
Blog
Explore the latest breakthroughs made possible with AI.
From deep learning model training and large-scale inference to enhancing operational efficiencies and customer experience, discover how AI is driving innovation and redefining the way organizations operate across industries. Source
Actions
Media Outlet details
| Scope | International |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesBuild Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills
AI agents are becoming a standard part of development workflows, but general-purpose agents weren’t built with specialized infrastructure software such as NVIDIA DOCA in mind. Without domain-specific knowledge, agents may fall back on guesswork. This is an issue in infrastructure development because every correction cycle takes time away from deployment.
Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples
Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems. Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux.
Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages
Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and local recording conditions are often underrepresented, so a multilingual model that performs well on broad benchmarks may still fall short in deployment. Saudi Arabic makes that concrete. A model may recognize Modern Standard Arabic or English yet struggle with Najdi and Hijazi speech, or local recording conditions.
Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton
Generative recommender (GR) systems are emerging as a powerful new approach for large-scale personalization. Instead of treating recommendation as a set of isolated retrieval, ranking, and prediction stages, GRs reformulate recommendation as sequence modeling over user behavior. A user’s interactions, context, candidate items, and actions become tokens in a high-cardinality event stream, and the model learns to generate or score the next relevant items from that sequence.
Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
AI infrastructure engineers, storage developers, and cloud service providers need fast and secure access to high-capacity file and object storage to support AI workloads. AI workloads increasingly require high-speed data access for training, fine-tuning, inference context, tool calls, searches, and database lookups. Much of this data lies in files and objects stored both on-premises and in the cloud.
Tracing Agent Harness Behavior with NVIDIA NeMo Relay
An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens. Inefficiencies create more chances for failure. To improve an agent’s behavior, developers must understand whether a task succeeded and how the agent completed it.
Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3
Vision-language models have made it possible to build visual AI agents that understand video at production scale. The harder problem is turning that capability into a maintainable system that combines ingestion, stream processing, event detection, retrieval, summarization, and reporting. The NVIDIA Metropolis Blueprint for Video Search and Summarization (VSS) and its agent skills help developers build visual AI agents faster.
AI Native by Design: Lessons Learned from Building NVIDIA TensorRT Model Connect
Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT Model Connect is an open source collection of AI model reference implementations in C++, built on top of NVIDIA TensorRT. It began with a practical question: could the performance of the NVIDIA inference stack be made accessible to model developers who are not TensorRT experts?
NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring
To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of possibilities. You could build a website over a weekend and share it with the world, or chat with someone half way around the world in online chat rooms without long-distance telephone fees. It brought endless opportunity, but also a lot of risk.
Add Runtime Controls to AI Agents with NVIDIA OpenShell
AI agents can be given a goal, write code, use tools, and keep working as new information becomes available. This opens the door to applications that investigate software failures, run experiments, and carry out business-critical actions and research over days or weeks. Useful agents need access to workspaces, compute resources, data, credentials, and external services.