How NVIDIA Uses Hugging Face
242 engineering articles about Hugging Face from NVIDIA's engineering team
Other NVIDIA Technologies
Other Companies Using Hugging Face
Articles
Filter:
Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems.
Luca Spindler
3 min read
Includes Code
--
Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data.
Elizabeth Goodman
13 min read
Includes Code
--
Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT…
Tanya Lenz
10 min read
Includes Code
--
As language models grow, scaling dense architectures becomes increasingly expensive. In a dense transformer, every token passes through every layer…
Michelle Horton
6 min read
Includes Code
--
Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and…
Tanya Lenz
11 min read
Includes Code
--
An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software…
Elizabeth Goodman
7 min read
Includes Code
--
When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment…
Elizabeth Goodman
10 min read
Includes Code
--
How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5…
Elizabeth Goodman
8 min read
--
NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two…
Elizabeth Goodman
11 min read
--
The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently…
Elizabeth Goodman
12 min read
Includes Code
--
A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle…
Michelle Horton
14 min read
Includes Code
--
Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect
Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing…
Tanya Lenz
6 min read
Includes Code
--
Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to…
Tanya Lenz
15 min read
Includes Code
--
Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.
Michelle Horton
3 min read
--
Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for…
Michelle Horton
10 min read
Includes Code
--
Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models…
Tanya Lenz
18 min read
Includes Code
--
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem.
Michelle Horton
4 min read
--
Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning…
Tanya Lenz
8 min read
--
Learn how NVIDIA NeMo Switchyard routes AI agent workloads across models using tuning-free and tunable routers that balance model capability, cost, and latency.
Michelle Horton
11 min read
--
Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI agentic…
Michelle Horton
4 min read
--
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene…
Michelle Horton
7 min read
--
Autonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding…
Elizabeth Goodman
12 min read
Includes Code
--
Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation.
Michelle Horton
11 min read
Includes Code
--
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they…
Tanya Lenz
4 min read
--
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…
Elizabeth Goodman
8 min read
--
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a…
Michelle Horton
12 min read
Includes Code
--
Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking…
Elizabeth Goodman
11 min read
Includes Code
--
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
Tanya Lenz
14 min read
Includes Code
--
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
Tanya Lenz
11 min read
Includes Code
--
Fine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings…
Elizabeth Goodman
13 min read
Includes Code
--
As more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing.
Elizabeth Goodman
10 min read
Includes Code
--
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context…
Tanya Lenz
11 min read
--
As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization…
Michelle Horton
15 min read
Includes Code
--
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.
Amr Elmeleegy
7 min read
--
Foundation models are reshaping computational biology. Pretrained on massive corpora of protein or genomic sequences, models such as ESM2 (a protein language…
Bruno Alvisio
11 min read
Includes Code
--
Developers building real-time AI—such as chat assistants, copilots, and agentic workflows—are often constrained by token-by-token generation speed.
Anu Srivastava
4 min read
Includes Code
--
As enterprise AI adoption scales, developers are increasingly forced to stitch together fragmented pipelines—separate models for text, vision…
Anu Srivastava
5 min read
Includes Code
--
Single-turn chatbots are evolving into long-running agents that can reason, maintain context, use tools, and run efficiently across many turns to complete…
Developing autonomous vehicle (AV) policies requires bridging an important gap between training and deployment. Vision-language-action (VLA) models that can…
Boris Ivanovic
8 min read
Includes Code
--
Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what’s…
Asawaree Bhide
11 min read
Includes Code
--
AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video…
Anu Srivastava
4 min read
Includes Code
--
Large language models (LLMs) are revolutionizing the financial trading landscape by enabling sophisticated analysis of vast amounts of unstructured data to…
Dan Blanaru
10 min read
Includes Code
--
This post is the second of a three-part series. See also Model Quantization: Concepts, Methods, and Why It Matters and Model Quantization: Turn FP8 Checkpoints…
Ruixiang Wang
8 min read
Includes Code
--
Agentic systems often reason across screens, documents, audio, video, and text within a single perception‑to‑action loop. However, they still rely on fragmented…
Anjali Shah
11 min read
--
DeepSeek just launched its fourth generation of flagship models with DeepSeek-V4-Pro and DeepSeek-V4-Flash, both targeted at enabling highly efficient million…
Anu Srivastava
6 min read
Includes Code
--
NVIDIA Ising is the world’s first family of open AI models for building quantum processors, launching with two model domains: Ising Calibration and Ising…
Tom Lubowe
9 min read
--
The release of MiniMax M2.7 adds enhancements to the popular MiniMax M2.5 model, built for agentic harnesses, and other complex use cases in fields such as…
Anu Srivastava
4 min read
Includes Code
--
In vision AI systems, model throughput continues to improve. The surrounding pipeline stages must keep pace, including decode, preprocessing, and GPU scheduling.
Andreas Kieslinger
9 min read
--
The Gemmaverse expands with the launch of the latest Gemma 4 multimodal and multilingual models, designed to scale across the full spectrum of deployments…
Anu Srivastava
6 min read
--
Developing new protein-based therapies and catalysts involves the challenging task of designing protein binders, or proteins that bind to a target protein or…
Kyle Gion
10 min read
Includes Code
--