NVIDIA logo

NVIDIA Engineering Blog & Tech Articles

Computing platform company pioneering GPU technology, AI infrastructure, and accelerated computing solutions for developers and data scientists

4696 engineering articles, tutorials, and technical insights from NVIDIA's engineering team

Latest Articles

Filter:
NVIDIA logo
NVIDIA
Intermediate
Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models…
Tanya Lenz
18 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem.
Michelle Horton
4 min read
--
NVIDIA logo
NVIDIA
Intermediate
AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades…
Jorge Cardoso
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare…
Elizabeth Goodman
7 min read
--
NVIDIA logo
NVIDIA
Intermediate
Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning…
NVIDIA logo
NVIDIA
Advanced
Learn how NVIDIA NeMo Switchyard routes AI agent workloads across models using tuning-free and tunable routers that balance model capability, cost, and latency.
NVIDIA logo
NVIDIA
Intermediate
Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI agentic…
Michelle Horton
4 min read
--
NVIDIA logo
NVIDIA
Intermediate
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene…
Michelle Horton
7 min read
--
NVIDIA logo
NVIDIA
Advanced
Autonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding…
Elizabeth Goodman
12 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Running a dedicated Kubernetes cluster per team often results in more isolation than an organization requires. While one cluster can be successfully shared…
Tanya Lenz
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data…
Elizabeth Goodman
12 min read
--
NVIDIA logo
NVIDIA
Advanced
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).
Tanya Lenz
13 min read
--
NVIDIA logo
NVIDIA
Advanced
The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration…
Elizabeth Goodman
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users…
Michelle Horton
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Beginner
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example…
Michelle Horton
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.
Elizabeth Goodman
12 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source…
Tanya Lenz
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation.
Michelle Horton
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they…
Tanya Lenz
4 min read
--
NVIDIA logo
NVIDIA
Advanced
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context…
Michelle Horton
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…
Elizabeth Goodman
8 min read
--
NVIDIA logo
NVIDIA
Advanced
As AI workloads increase, explosive compute demand is pushing the semiconductor industry to meet unprecedented performance targets. Even small delays can have…
Tanya Lenz
6 min read
--
NVIDIA logo
NVIDIA
Intermediate
Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse…
Elizabeth Goodman
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways…
Tanya Lenz
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a…
Michelle Horton
12 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can…
Michelle Horton
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.
Tanya Lenz
13 min read
--
NVIDIA logo
NVIDIA
Advanced
Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context…
Michelle Horton
12 min read
--
NVIDIA logo
NVIDIA
Advanced
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token…
Kirthi Devleker
7 min read
--
NVIDIA logo
NVIDIA
Advanced
The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute…
Elizabeth Goodman
13 min read
--
NVIDIA logo
NVIDIA
Intermediate
Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and…
Tanya Lenz
16 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Capcom’s RE ENGINE team set out to bring path tracing into two shipping titles at once, Resident Evil Requiem and PRAGMATA, each with a different visual…
Michelle Horton
8 min read
--
NVIDIA logo
NVIDIA
Advanced
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and…
Tanya Lenz
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks…
Michelle Horton
11 min read
--
NVIDIA logo
NVIDIA
Advanced
Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking…
Elizabeth Goodman
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data…
Michelle Horton
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
For over fifteen years, x86 CPUs have shipped with a dedicated hardware instruction for carryless multiplication. It’s a small but stubborn primitive that sits…
Michelle Horton
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when…
Elizabeth Goodman
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
Tanya Lenz
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
NVIDIA logo
NVIDIA
Advanced
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes…
Elizabeth Goodman
5 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Across science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these…
Elizabeth Goodman
7 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Robotics foundation models have made remarkable progress. Today’s best systems can follow natural language instructions to pick, place, sort…
Brad Nemire
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states…
Tanya Lenz
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead…
Michelle Horton
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means…
Elizabeth Goodman
16 min read
--
NVIDIA logo
NVIDIA
Advanced
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design.
Elizabeth Goodman
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings…
Elizabeth Goodman
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Molecular dynamics (MD) simulations are among the most demanding workloads in computational science. Using them, researchers can observe atomic behavior in…
Michelle Horton
21 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are…
Sean Lopp
10 min read
Includes Code
--