NVIDIA logo

How NVIDIA Uses Large Language Models

40 engineering articles about Large Language Models from NVIDIA's engineering team

Articles

Filter:
NVIDIA logo
NVIDIA
Advanced
AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means…
Elizabeth Goodman
16 min read
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the NVIDIA Nemotron 3, a family of open models designed for agentic AI systems, emphasizing its efficiency and accuracy through innovative architectures and techniques.
NVIDIA logo
NVIDIA
Intermediate
The NVIDIA Blackwell architecture has achieved the fastest training times across all MLPerf Training v5. 1 benchmarks, showcasing significant advancements in AI training performance.
NVIDIA logo
NVIDIA
Advanced
The article discusses how to enhance the efficiency of Large Language Models (LLMs) during inference by utilizing CPU-GPU memory sharing through NVIDIA's NVLink C2C technology.
Afroze Syed
6 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
This article serves as a comprehensive guide for benchmarking Large Language Models (LLMs) using NVIDIA's GenAI-Perf tool alongside NVIDIA NIM.
Vinh Nguyen
11 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
This article provides a comprehensive guide on Vision Language Models (VLMs) and their evolution from single-image understanding to advanced video comprehension.
Shubham Agrawal
11 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the Marco framework, a configurable graph-based task-solving and multi-AI agent system designed to streamline chip design processes.
Mark Ren
8 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA has released a new Generative AI Teaching Kit aimed at enhancing education in generative AI technologies.
NVIDIA logo
NVIDIA
Intermediate
The article discusses the implementation of data-efficient knowledge distillation using NVIDIA NeMo-Aligner during supervised fine-tuning (SFT).
Anna Shors
5 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA showcased its AI security expertise at the Black Hat USA and DEF CON conferences, focusing on the evolving landscape of AI in cybersecurity.
NVIDIA logo
NVIDIA
Intermediate
The article discusses the deployment of diverse AI applications using Multi-LoRA support on NVIDIA RTX AI PCs and workstations.
Annamalai Chockalingam
9 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses how NVIDIA NVLink and NVSwitch enhance the performance of Large Language Model (LLM) inference by enabling efficient multi-GPU computing.
Brian Slechta
7 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the launch of Continuum AI by Edgeless Systems, a generative AI framework that ensures data privacy through confidential computing and NVIDIA H100 GPUs.
Laura Martinez
6 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
This article explores the complexities of deploying trillion-parameter large language models (LLMs) in production environments, focusing on maximizing throughput and user interactivity.
Amr Elmeleegy
13 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the deployment of LoRA (Low-Rank Adaptation) fine-tuned models using NVIDIA NIM, highlighting the advantages of customizing large language models (LLMs) for specific tasks.
Shashank Verma
11 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article provides a comprehensive guide on deploying generative AI using NVIDIA NIM microservices, highlighting its ease of use for enterprise developers in both on-premises and cloud environmen...
NVIDIA logo
NVIDIA
Advanced
The article discusses the integration of AI chatbots, particularly Gipi, with NVIDIA TensorRT-LLM and AI foundation models to enhance personalized learning experiences.
Nisanur Genc
5 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses how Union. ai and NVIDIA DGX Cloud are transforming AI workflows by providing accessible, high-performance computing resources.
NVIDIA logo
NVIDIA
Advanced
The article discusses the Low-Rank Adaptation (LoRA) method for fine-tuning large language models (LLMs) using NVIDIA TensorRT-LLM.
Amit Bleiweiss
15 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the upcoming NVIDIA GTC 2024 event, highlighting the benefits of in-person attendance, including networking opportunities, hands-on training, and exclusive sessions focused on...
NVIDIA logo
NVIDIA
Intermediate
The article discusses the evaluation of Retrieval-Augmented Generation (RAG) systems, emphasizing the importance of embedding models and systematic evaluation processes.
Benedikt Schifferer
14 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article reviews the most popular NVIDIA Technical Blog posts of 2023, highlighting advancements in generative AI, large language models (LLMs), high-performance computing (HPC), and robotics.
NVIDIA logo
NVIDIA
Intermediate
The article discusses how AI-powered note-taking and summarization can enhance meeting productivity by leveraging a cloud-native microservice architecture.
Mohamed Elshenawy
6 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the rising threat of identity-based attacks, particularly phishing, and how Generative AI and Large Language Models (LLMs) are transforming cybersecurity.
NVIDIA logo
NVIDIA
Intermediate
The article discusses the intricacies of training Large Language Models (LLMs) using transformer networks, focusing on model architectures, attention mechanisms, and embedding techniques.
NVIDIA logo
NVIDIA
Intermediate
The article discusses deploying large language models (LLMs) at the edge using the NVIDIA IGX Orin Developer Kit.
NVIDIA logo
NVIDIA
Advanced
The article discusses how NVIDIA's H100 GPUs and Quantum-2 InfiniBand have set new performance records in data center-scale AI training, particularly for Large Language Models (LLMs) and Stable Dif...
NVIDIA logo
NVIDIA
Intermediate
The article discusses the application of Large Language Models (LLMs) in enterprise solutions, highlighting their capabilities in enhancing productivity across various industries.
NVIDIA logo
NVIDIA
Advanced
NVIDIA has released TensorRT-LLM, an open-source library designed to optimize inference performance for large language models (LLMs) on NVIDIA GPUs.
Neal Vaidya
10 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses various techniques for customizing Large Language Models (LLMs) to better fit enterprise needs, emphasizing the importance of tailoring language processing capabilities for sp...
NVIDIA logo
NVIDIA
Intermediate
The article discusses strategies for improving outputs from Large Language Models (LLMs) by focusing on prompt design and parameter tuning.
Annie Surla
12 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
This article provides an introduction to Large Language Models (LLMs), focusing on prompt engineering and P-tuning techniques.
NVIDIA logo
NVIDIA
Intermediate
NVIDIA has introduced generative AI services aimed at enhancing language, visual content, and biology applications.
NVIDIA logo
NVIDIA
Intermediate
The article discusses NVIDIA's BioNeMo service, a framework for training and serving biomolecular large language models (LLMs) designed for predicting protein structures and properties.
Vanessa Braunstein
3 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the advancements in AI models, particularly NVIDIA's pretrained models, which have significantly impacted various industries in 2022.
Pranjali Joshi
7 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the latest SDKs available in the NGC catalog, focusing on tools for Large Language Models (LLMs), digital twins, and digital biology.
NVIDIA logo
NVIDIA
Intermediate
The article discusses NVIDIA's efforts to simplify access to large language models (LLMs) through the NeMo framework and associated services, including NeMo LLM and BioNeMo.
Annamalai Chockalingam
4 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA has announced significant updates to the NeMo framework, enhancing the training speed of large language models (LLMs) by up to 30%.
NVIDIA logo
NVIDIA
Intermediate
The article discusses major updates to NVIDIA's Riva SDK for building speech AI applications and the NeMo framework for training large language models.
Siddharth Sharma
3 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
At GTC 2022, NVIDIA unveiled significant updates to its AI software suite, focusing on advancements in speech AI, recommenders, and inference optimization. The updates include the launch of Riva 2.
Siddharth Sharma
5 min read
Has Summary
--

You've reached the end! All 40 articles loaded.