NVIDIA logo

How NVIDIA Uses PyTorch

602 engineering articles about PyTorch from NVIDIA's engineering team

Articles

Filter:
NVIDIA logo
NVIDIA
Advanced
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem.
Michelle Horton
4 min read
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users…
Michelle Horton
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token…
Kirthi Devleker
7 min read
--
NVIDIA logo
NVIDIA
Intermediate
Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and…
Tanya Lenz
16 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes…
Elizabeth Goodman
5 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead…
Michelle Horton
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design.
Elizabeth Goodman
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high-fidelity 3D representations of real-world environments from multisensor data such…
Tanya Lenz
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines…
Peter Kisfaludi
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
Amr Elmeleegy
7 min read
--
NVIDIA logo
NVIDIA
Advanced
Foundation models are reshaping computational biology. Pretrained on massive corpora of protein or genomic sequences, models such as ESM2 (a protein language…
Bruno Alvisio
11 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
As enterprise AI adoption scales, developers are increasingly forced to stitch together fragmented pipelines—separate models for text, vision…
Anu Srivastava
5 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
This post is the third of a three-part series. See also Model Quantization: Concepts, Methods, and Why It Matters and Model Quantization: Post-Training…
Ruixiang Wang
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
AI agents are changing how you interact with your PC. Creators, developers, and AI enthusiasts are already using these agents extensively to assist with day-to…
Annamalai Chockalingam
8 min read
--
NVIDIA logo
NVIDIA
Advanced
Each wave of AI has created a new scaling law. Pretraining scaled intelligence through larger datasets, more parameters, and massively parallel GPU systems.
Praveen Menon
7 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video…
Anu Srivastava
4 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
In production inference deployments, demand fluctuates over time, requiring inference replicas to scale elastically. However, cold-starting inference workloads…
Schwinn Saereesitthipitak
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Large language models (LLMs) are revolutionizing the financial trading landscape by enabling sophisticated analysis of vast amounts of unstructured data to…
Dan Blanaru
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA CUDA 13.3 brings new capabilities and performance optimizations to developers across the CUDA ecosystem. The launch of NVIDIA CUDA Tile programming in…
Jonathan Bentz
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
The path from a trained AI model to production should be smooth, but rarely is. Many teams invest weeks fine-tuning models, only to discover that exporting to a…
Lovina Dmello
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
This post is the second of a three-part series. See also Model Quantization: Concepts, Methods, and Why It Matters and Model Quantization: Turn FP8 Checkpoints…
Ruixiang Wang
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
For decades, computational biology has operated under a reductionist compromise. To fit complex biological systems into the limited memory of a single GPU…
Dejun Lin
8 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Federated learning (FL) is no longer a research curiosity—it’s a practical response to a hard constraint: the most valuable data is often the least movable.
NVIDIA logo
NVIDIA
Intermediate
In March 2026, three LLM agents generated over 600,000 lines of code, ran 850 experiments, and helped secure a first-place finish in a Kaggle playground…
NVIDIA logo
NVIDIA
Advanced
In a previous post, we introduced the Universal Sparse Tensor (UST), enabling developers to decouple a tensor’s sparsity from its memory layout for greater…
Aart J.C. Bik
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Higher-order optimization algorithms such as Shampoo have been effectively applied in neural network training for at least a decade. These methods have achieved…
Hao Wu
9 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
The development of socially acceptable nuclear reactors requires that they are safe, clean, efficient, economical, and sustainable.
Mark Hobbs
11 min read
--
NVIDIA logo
NVIDIA
Advanced
For decades, computational chemistry has faced a tug-of-war between accuracy and speed. Ab initio methods like density functional theory (DFT) provide high…
Erica Tsai
13 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA Ising is the world’s first family of open AI models for building quantum processors, launching with two model domains: Ising Calibration and Ising…
NVIDIA logo
NVIDIA
Intermediate
Training LLMs requires periodic checkpoints. These full snapshots of model weights, optimizer states, and gradients are saved to storage so training can resume…
Wenqi Glantz
12 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Physical AI—AI systems that perceive, reason, and act in physically grounded simulated environments—is changing how teams design and validate robots and…
NVIDIA logo
NVIDIA
Intermediate
The Gemmaverse expands with the launch of the latest Gemma 4 multimodal and multilingual models, designed to scale across the full spectrum of deployments…
NVIDIA logo
NVIDIA
Advanced
NVIDIA Groq 3 LPX is a new rack-scale inference accelerator for the NVIDIA Vera Rubin platform, designed for the low-latency and large-context demands of…
NVIDIA logo
NVIDIA
Advanced
Computer-aided engineering (CAE) is shifting from human-driven workflows toward AI-driven ones, including physics foundation models that generalize across…
NVIDIA logo
NVIDIA
Intermediate
CUDA 13.2 arrives with a major update: NVIDIA CUDA Tile is now supported on devices of compute capability 8.X architectures (NVIDIA Ampere and NVIDIA Ada)…
Jonathan Bentz
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Alibaba has introduced the new open source Qwen3.5 series built for native multimodal agents. The first model in this series is a ~400B parameter native vision…
Anu Srivastava
3 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the use of NVFP4 low-precision model training to achieve higher throughput without sacrificing accuracy in AI model training.
Aditya Vavre
7 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses how the NVIDIA cuda. compute library enables Python developers to write high-performance GPU code without needing to resort to C++.
Daniel Rodriguez
5 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses how NVIDIA's hardware-software co-design significantly enhanced the inference performance of Sarvam AI's Sovereign 30B model, achieving a 4x speedup on NVIDIA Blackwell archit...
Utkarsh Uppal
14 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses NVIDIA TensorRT LLM AutoDeploy, a beta feature that automates the inference optimization process for large language models (LLMs).
​​Lucas Liebenwein
8 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
Kimi K2. 5 is an advanced multimodal vision language model (VLM) developed by Kimi, optimized for various AI tasks.
Anu Srivastava
4 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the challenges of Expert Parallel communication in training Mixture-of-Experts (MoE) models and introduces Hybrid-EP, an efficient communication solution that leverages NVIDIA...
Fan Yu
10 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the Universal Sparse Tensor (UST), a framework designed to efficiently handle sparse tensors across various applications, including scientific computing and deep learning.
Aart J.C. Bik
13 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the transition from the traditional two-phase API of the CUB library to a new single-call API introduced in CUDA 13. 1.
Giannis Gonidelis
8 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
This article provides a detailed guide on implementing high-performance matrix multiplication using NVIDIA's cuTile framework in CUDA.
Jinman Xie
13 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses NVIDIA's advancements in AI model inference performance through the Blackwell architecture, emphasizing improvements in token throughput per watt and the enhancements made to ...
Ashraf Eassa
5 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses how recent upgrades to open source AI tools enhance the performance of small language models (SLMs) and diffusion models on NVIDIA RTX PCs.
Annamalai Chockalingam
7 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the latest software and model optimizations for NVIDIA DGX Spark, highlighting significant performance improvements in AI workflows.
Allen Bourgoyne
5 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the NVIDIA Rubin platform, which introduces six new chips designed to create a powerful AI supercomputer.
NVIDIA logo
NVIDIA
Advanced
NVIDIA introduces the Jetson T4000, enhancing AI and real-time reasoning for robotics and edge AI applications with up to 1200 FP4 TFLOPs of AI compute and 64 GB of memory.
Shashank Maheshwari
9 min read
Has Summary
--