#

Reinforcement Learning Programming Tutorials & Engineering Articles

85 Reinforcement Learning tutorials, guides, and engineering insights from NVIDIA, Uber, Netflix, and more

Reinforcement Learning Articles & Tutorials

Filter:
Shopify logo
Shopify
Advanced
How we compress production failures into model weights every day, beat frontier-model quality, and cut serving costs 96%.
NVIDIA logo
NVIDIA
Advanced
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a…
Michelle Horton
12 min read
Includes Code
--
Google logo
Google
Advanced
Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurrent, asynchronous rollouts with a decoupled producer-consumer pipeline, ensuring the trainer is constantly fed even while agents wait on network I/O or environment steps. Additionally, Tunix provides plug-and-play abstractions and continuous macro-level profiling, allowing developers to easily integrate custom open-source environments and optimize complex distributed workflows without massive code rewrites.
Haoyu Gao, Lance Wang, Shadi Noghabi, Tianshu Bao, Weiren Yu
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Robotics foundation models have made remarkable progress. Today’s best systems can follow natural language instructions to pick, place, sort…
Brad Nemire
14 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer…
Elizabeth Goodman
12 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Single-turn chatbots are evolving into long-running agents that can reason, maintain context, use tools, and run efficiently across many turns to complete…
Google logo
Google
Advanced
The Google Tunix Hackathon on Kaggle challenged developers to transform small, non-reasoning base models into general reasoning engines using Kaggle TPUs and a limited compute budget. The winning teams achieved this by implementing multi-stage post-training pipelines that combined Supervised Fine-Tuning (SFT) with advanced alignment techniques like GRPO and SimPO. Ultimately, the competition democratized AI development by proving that highly capable, structured reasoning models can be successfully trained by the community using accessible, open-source resources.
Wei Wei, Weiren Yu, Tianshu Bao, Lance Wang, Chris Achard
6 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
As LLMs transition from simple text generation to complex reasoning, reinforcement learning (RL) plays a central role. Algorithms like Group Relative Policy…
Guyue Huang
8 min read
Includes Code
--
Google logo
Google
Intermediate
MaxText has introduced new support for Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on single-host TPU configurations, leveraging JAX and the Tunix library for high-performance model refinement. These features enable developers to easily adapt pre-trained models for specialized tasks and complex reasoning using efficient algorithms like GRPO and GSPO. This update streamlines the post-training workflow, offering a scalable path from single-host setups to larger multi-host configurations.
Wei Wei, Weiren Yu
3 min read
Includes Code
--
Meta logo
Meta
Advanced
This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation. The previous post introduced Ranking Eng…
NVIDIA logo
NVIDIA
Advanced
This article explores how to train an AI agent to operate a new Command Line Interface (CLI) using synthetic data generation and reinforcement learning.
Chris Alexiuk
11 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the NVIDIA Nemotron 3, a family of open models designed for agentic AI systems, emphasizing its efficiency and accuracy through innovative architectures and techniques.
NVIDIA logo
NVIDIA
Intermediate
The article discusses the development of scientific AI agents using reinforcement learning (RL) techniques, specifically through the NVIDIA NeMo framework.
Christian Munley
12 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article introduces Broadened Reinforcement Learning (BroRL), a new paradigm that enhances the training of large language models (LLMs) by focusing on rollout scaling rather than just increasing...
Jian Hu
6 min read
Has Summary
--
Netflix logo
Netflix
Advanced
Netflix introduces Advantage-Weighted Supervised Fine-Tuning (A-SFT), a novel post-training algorithm for generative recommender systems that addresses the unique challenges of applying reinforceme...
Netflix Technology Blog
12 min read
Has Summary
--
Google logo
Google
Advanced
The article introduces Tunix, a new open-source, JAX-native library designed for post-training of large language models (LLMs).
Srikanth Kilaru, Tianshu Bao
7 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
This article discusses the integration of the Newton physics engine with NVIDIA Isaac Lab for training quadruped locomotion policies and simulating cloth manipulation.
Mohammad Mohajerani
13 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the enhancements in reinforcement learning training throughput using NVIDIA NeMo-RL with Megatron-Core support.
Anna Shors
7 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the advancements in reinforcement learning for large language models (LLMs) through the introduction of ProRL v2 by NVIDIA Research.
Jian Hu
7 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the release of the NVIDIA Llama Nemotron Super 49B v1. 5, highlighting its advancements in accuracy, efficiency, and reasoning capabilities for AI agents.
Chris Alexiuk
5 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article introduces NVIDIA NeMo-RL, an open-source library for reinforcement learning that supports scalable training from single-GPU to thousand-GPU models.
Alexander Bukharin
5 min read
Includes Code
Has Summary
--
Uber logo
Uber
Advanced
This article discusses how Uber utilizes reinforcement learning techniques to enhance the efficiency of its marketplace by improving the balance between drivers and demand.
Prateek Jain, Soheil Sadeghi, Mehrdad Bakhtiari
11 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the advancements in AI autonomy through NVIDIA's Nemotron open reasoning models, which enhance AI agents' decision-making capabilities in complex environments.
Nirmal Kumar Juluru
6 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the development and capabilities of NVIDIA's Llama Nemotron reasoning models, which enhance AI agents' reasoning abilities for complex problem-solving in various industries.
Chris Alexiuk
11 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses how NVIDIA Cosmos World Foundation Models (WFMs) enhance the development of AI-driven robots and autonomous vehicles by providing high-fidelity, physics-aware synthetic data.
Pranjali Joshi
7 min read
Includes Code
Has Summary
--
Google logo
Google
Intermediate
Gemma 3 is the latest version of the Gemma open-model family, boasting enhanced capabilities such as multimodality, longer context windows, and improved reasoning.
Omar Sanseviero, Philipp Schmid
5 min read
Includes Code
Has Summary
--
LinkedIn logo
LinkedIn
Advanced
The article discusses the development of domain-adapted foundation GenAI models at LinkedIn, focusing on their application within the Economic Opportunity Network (EON) project.
Airbnb logo
Airbnb
Advanced
The article discusses Airbnb's evolution in location retrieval from simple heuristics to advanced machine learning and reinforcement learning techniques.
Dillon Davis
9 min read
Has Summary
--
Netflix logo
Netflix
Intermediate
The article discusses Netflix's approach to enhancing long-term member satisfaction through its recommendation algorithms.
Netflix Technology Blog
9 min read
Has Summary
--
Pinterest logo
Pinterest
Intermediate
The article discusses the development of Pinterest Canvas, a text-to-image foundation model aimed at enhancing existing images and products on the Pinterest platform.
NVIDIA logo
NVIDIA
Advanced
The article discusses the challenges and methodologies involved in training quadruped locomotion policies using NVIDIA Isaac Lab, emphasizing the importance of high-fidelity simulation for bridging...
Oyindamola Omotuyi
11 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the advancements in robotics workflows through the latest release of NVIDIA Isaac Sim 4. 0 and NVIDIA Isaac Lab.
Akhil Docca
10 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses the integration of AI in enabling practical quantum computing by addressing challenges in quantum processors, error correction, and algorithm development.
Mark Wolf
6 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article introduces DRaFT+, an enhanced algorithm for fine-tuning text-to-image diffusion models, which aims to improve the alignment between input prompts and generated images.
Ali Taghibakhshi
9 min read
Includes Code
Has Summary
--
Fly.io logo
Fly.io
Intermediate
The article discusses Yoko Li's innovative work in AI, focusing on her projects like AI Town and AI Tamago, which utilize emergent behavior and large language models.
Palantir logo
Palantir
Intermediate
The article discusses how to leverage Palantir AIP to build a semantic search application that uncovers insights from unstructured data within enterprises.
NVIDIA logo
NVIDIA
Advanced
The article discusses how NVIDIA NeMo can streamline the development of generative AI applications on GPU-accelerated Google Cloud.
NVIDIA logo
NVIDIA
Advanced
QHack 2023 showcased the intersection of quantum computing and machine learning, featuring 2,850 participants from 105 countries competing to develop innovative solutions using NVIDIA's quantum tec...
Tom Lubowe
8 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses how AutoDMP leverages AI and GPU technology to optimize macro placement in chip design, significantly improving performance and efficiency.
Anthony Agnesina
10 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article summarizes the advancements and AI-powered solutions introduced in 2022, highlighting the most popular posts on the NVIDIA Technical Blog.
Michelle Horton
3 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Beginner
The article discusses the DeXtreme project, which utilizes simulation to teach dexterity to a real robot hand.
Gavriel State
7 min read
Has Summary
--
Netflix logo
Netflix
Advanced
This article discusses the application of reinforcement learning to create optimal recommendation systems that consider users' time budgets.
Netflix Technology Blog
13 min read
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA researchers introduced Factory, a novel simulation approach designed to enhance robotic assembly by enabling real-time, accurate simulations of contact-rich interactions.
Oyindamola Omotuyi
10 min read
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
The article discusses the innovative use of deep reinforcement learning (RL) to design arithmetic circuits, particularly in the context of NVIDIA GPUs.
Rajarshi Roy
8 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Advanced
The article discusses NVIDIA's advancements in MLPerf Training v2. 0, highlighting the full-stack optimizations that enhance performance across various AI workloads.
NVIDIA logo
NVIDIA
Advanced
The article discusses the performance improvements achieved in the NVIDIA MLPerf Training v1. 1 benchmark through full stack optimization.
Vinh Nguyen
21 min read
Includes Code
Has Summary
--
NVIDIA logo
NVIDIA
Intermediate
NVIDIA has released an updated Edge AI and Robotics Teaching Kit aimed at university educators, developed in collaboration with experts from the University of Oxford and the University of Maryland,...
NVIDIA logo
NVIDIA
Advanced
The article discusses NVIDIA's research on transferring dexterous manipulation capabilities from GPU simulation to real-world robotic applications.
Varun Lodaya
11 min read
Has Summary
--
Airbnb logo
Airbnb
Advanced
This article discusses how Airbnb utilizes task-oriented conversational AI to enhance customer support for hosts and guests.