#

Apache Programming Tutorials & Engineering Articles

1019 Apache tutorials, guides, and engineering insights from LinkedIn, Uber, NVIDIA, and more

Apache Articles & Tutorials

Filter:
NVIDIA logo
NVIDIA
Intermediate
To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of…
Tanya Lenz
7 min read
--
NVIDIA logo
NVIDIA
Intermediate
Using NVIDIA Cluster Readiness Engine, teams can bring reliable GPU clusters to production with workload-driven validation.
Michelle Horton
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Advanced
Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents…
Michelle Horton
10 min read
Includes Code
--
Paarth Chothani, Chirag Agrawal
11 min read
--
Meta logo
Meta
Advanced
We’re open-sourcing Rebalancer, the assignment-problem solver that has been used to solve resource allocation problems throughout Meta for over nine years. Rebalancer separates several relate…
Richard Barnes
10 min read
--
ClickHouse logo
ClickHouse
Beginner
Notes from three days in the Netherlands, featuring a lightning talk on pg_clickhouse and pg_stat_ch at PGDay Lowlands and a session on PostgreSQL 19 monitoring at Percona Live Amsterdam.
8 min read
Includes Code
--
ClickHouse logo
ClickHouse
Intermediate
The official ClickHouse provider for Apache Airflow simplifies data workflows with standard SQL operators, bulk inserts, and shared setup across self-managed Airflow and Astronomer.
7 min read
Includes Code
--
ClickHouse logo
ClickHouse
Intermediate
The ClickHouse adapter for dbt v2 is in public beta, powered by dbt’s Rust engine. ClickHouse also joins the dbt platform in private beta, supporting open-source ClickHouse and ClickHouse Cloud.
10 min read
Includes Code
--
Google logo
Google
Intermediate
Google has partnered with Speakeasy to open-source their OpenAPI code generation suite under the AGPLv3 license, a strategic move prompted by the sudden shutdown of Google's previous proprietary SDK provider. The newly open-sourced suite equips developers with deterministic, multi-language SDK generators that natively support strict typing and SSE streaming, alongside tools for compiling agent-native CLIs and documentation MCP servers. Engineering teams can now safely integrate this robust tooling directly into their CI pipelines to automatically generate reliable client libraries for their own APIs, all while retaining complete licensing control over the output code.
Amir Hardon, Philipp Schmid
3 min read
Includes Code
--
ClickHouse logo
ClickHouse
Advanced
ClickHouse 26.8 LTS introduces background queries, pipelined SQL, new text tokenizers, expanded data lake integrations, and faster Parquet, aggregation, and join queries.
ClickHouse logo
ClickHouse
Advanced
After claims that ClickHouse is “winning the observability wars” sparked debate, we reflect on why it has become a leading storage and query engine, where it still falls short, and why winning the database layer isn’t the same as winning observability.
NVIDIA logo
NVIDIA
Advanced
For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain…
Elizabeth Goodman
12 min read
Includes Code
--
Pankaj Mohapatra, Arun Mahadeva Iyer, Balajee Nagasubramaniam, Prashant Wason, Alon Levy
11 min read
--
Aravind Selvan, Calder Lund, and Gaurang Sadekar
13 min read
Includes Code
--
Nimish Sheth, Manas Kelshikar, Dhirendra Kumar Singh, Rajan Jana, Wasim Raza
8 min read
--
ClickHouse logo
ClickHouse
Intermediate
ClickHouse is joining the Open Secure AI Alliance alongside NVIDIA and other industry leaders to help build open tools that keep AI agents secure.
Google logo
Google
Advanced
To prevent context window bloat and reduce token consumption, Genkit Go introduces Agent Skills based on a progressive disclosure architecture. Developers can package specialized instructions, scripts, and references into modular SKILL.md bundles where only the frontmatter metadata is initially exposed to the agent's system prompt. When a task matches the skill's description, Genkit's middleware dynamically loads the full instruction body and associated assets, ensuring the model accesses precise workflows exactly when needed.
Prakhar Agarwal, Avinash Varma Sagi, Abhay Singh Chauhan, Kranthi Reddy, Divya Rai
11 min read
--
Ramp logo
Ramp
Beginner
Bypassing row-by-row Snowflake result materialization cut fetch-memory growth and doubled one ML workflow’s training-data window on the same cluster size.
5 min read
Includes Code
--
ClickHouse logo
ClickHouse
Advanced
Retention limits, sampling, and metric roll-ups aren't observability best practices - they're workarounds for storage systems that can't handle full-fidelity data, and they're becoming a hard blocker for AI-driven workflows.
19 min read
Includes Code
--
ClickHouse logo
ClickHouse
Beginner
Query Insights is now in preview for ClickHouse Cloud Managed Postgres: every query pattern your database runs, ranked by impact, with the diagnostic picture of why each one is slow.
9 min read
Includes Code
--
ClickHouse logo
ClickHouse
Intermediate
ClickHouse was released in open source on Jun 15 2016, ten years ago. Since then, it became the most popular open source analytical database with more than 2000 contributors.
ClickHouse logo
ClickHouse
Advanced
On June 15, 2016, ClickHouse went open source under the Apache 2.0 license. Ten years later, it's one of the most widely deployed analytical databases in the world.
ClickHouse logo
ClickHouse
Intermediate
ClickHouse now has an official ADBC driver, giving Ruby, R, C, and every other ADBC-aware tool zero-conversion, Arrow-native access to ClickHouse without a dedicated client for each language.
NVIDIA logo
NVIDIA
Advanced
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes…
Elizabeth Goodman
5 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
As more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing.
Elizabeth Goodman
10 min read
Includes Code
--
NVIDIA logo
NVIDIA
Intermediate
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context…
NVIDIA logo
NVIDIA
Advanced
GPU-accelerated query engines are often constrained by memory and I/O bandwidth. NVIDIA hardware advances—including high bandwidth memory (HBM)…
Michelle Horton
12 min read
--
Google logo
Google
Intermediate
An open specification for finding and verifying tools, skills, and agents across the web.Agents are ...
Junjie Bu, Srinivas Krishnan
5 min read
Includes Code
--
Google logo
Google
Intermediate
DiffusionGemma is an experimental text-generation model built on the Gemma 4 architecture that uses diffusion-based parallel generation instead of token-by-token autoregression, enabling much faster inference, bidirectional context awareness, and real-time self-correction while remaining deployable on consumer GPUs. Its architecture generates and refines 256-token blocks in parallel through iterative denoising, allowing it to handle complex constraint-based tasks such as Sudoku more effectively than traditional language models and demonstrating strong gains from fine-tuning. The model integrates with vLLM and other popular inference frameworks, giving developers access to a new non-autoregressive approach that combines high performance, efficient long-context scaling, and straightforward customization and deployment.
Ian Ballantyne, Omar Sanseviero
6 min read
Includes Code
--
Stripe logo
Stripe
Intermediate
Some microservices are difficult to test because their behavior depends on a long tail of inputs that are hard to model by hand. This post focuses on how to build a repeatable Apache Spark replay harness for testing without creating a second implementation of the service.
Vivek Yadav
8 min read
Includes Code
--
Netflix Technology Blog
13 min read
Includes Code
--
Stripe logo
Stripe
Intermediate
Some microservices are difficult to test because their behavior depends on a long tail of inputs that are hard to model by hand. For that class of problem, Apache Spark's massive parallelism and linear scalability can be leveraged to build a regression test harness. That makes it possible to compare old and new implementations, model upcoming rule changes, and quantify impact before a change reaches production.
Vivek Yadav
9 min read
--
NVIDIA logo
NVIDIA
Advanced
Maximizing the value of AI infrastructure demands deep visibility into GPU utilization. Yet many platform teams running AI workloads on Kubernetes operate with…
Guy Saltoun
6 min read
Includes Code
--
Airbnb logo
Airbnb
Advanced
How Airbnb shifts from PaaS to an internal knowledge graph infrastructure at scale.
Lucen Zhao
9 min read
--
NVIDIA logo
NVIDIA
Advanced
NVIDIA Ising is the world’s first family of open AI models for building quantum processors, launching with two model domains: Ising Calibration and Ising…