Skip to main content

52 docs tagged with "Kafka"

View all tags

Apache Kafka Architecture & Core Fundamentals

Complete architectural overview of Apache Kafka — origins, the M×N integration problem, the 4 pillars of Kafka's extreme speed (sequential I/O, page cache, zero-copy, batching), message anatomy, and the dumb broker / smart consumer paradigm.

Apache Kafka Knowledge Base

Apache Kafka is a distributed event streaming platform designed for high-throughput, fault-tolerant, and scalable real-time data pipelines and streaming applications.

Change Data Capture (CDC)

Comprehensive guide on Change Data Capture (CDC), detailing how it works, alternatives comparison, implementation patterns with Debezium and Spring, and deep dives for senior engineers.

Consumer Groups

A **consumer group** is a set of consumers that collectively consume a topic's partitions. Each partition is assigned to exactly one consumer within the group.

Consumer Lag, Poison Messages & Retry Topics

Deep dive into Kafka consumer lag — offset mechanics, lag calculation, root cause diagnosis, scaling strategies, and non-blocking retry topics (DLQ, multi-tier backoff, 2-phase consumer commits).

Dead Letter Queue (DLQ) Pattern

A comprehensive guide to the Dead Letter Queue (DLQ) pattern — covering poison pill handling, retry strategies, alternatives comparison, AWS SQS / Kafka / RabbitMQ implementations, and production deep dives for senior engineers.

Deployment Configuration & Infrastructure Verification

A comprehensive guide to managing and verifying application configurations, environment variables priority, HashiCorp Vault secrets, Kafka topics, ACLs, schema registry compatibility, database migrations, and post-deploy health verification.

Event-Driven Microservices

In-depth guide to Event-Driven Microservices, covering domain events, event types (notification vs carried-state vs sourcing), choreography sagas, transactional outbox pattern, schema evolution, ordering, idempotency, and Kafka setups.

Hash Key Partitions

Kafka uses a hash of the message key to determine partition assignment. Understanding this mechanism is essential for ordering guarantees, avoiding hot partitions, and designing correct partition keys.

Idempotent Producer

Without idempotence, network retries create duplicate messages on Kafka brokers. Enabling idempotence guarantees exactly-once delivery per producer session.

Kafka ACLs & Authorization Patterns

Kafka Access Control Lists (ACLs) for fine-grained authorization. Covers KRaft ACL storage, resource patterns, OAuth/OPA/RBAC integration, and at-scale management.

Kafka Broker — Complete Guide

A complete guide to Kafka brokers — what they are, how storage works, partition leadership, replication, ISR, KRaft vs ZooKeeper, log compaction, performance internals, and production monitoring. Beginner through senior depth.

Kafka Connect

**Kafka Connect** is a framework for **reliably moving data between Kafka and external systems** (databases, file systems, cloud services) without writing.

Kafka Consumer

A **consumer** reads messages from Kafka topics. Unlike traditional queues (push-based), Kafka consumers **pull** messages at their own pace. This gives.

Kafka Data Governance

The six primitives of Kafka data governance — schema policy, topic ownership, access control, encryption/masking, audit/lineage, and data quality. Why brokers don't provide governance and how to build it.

Kafka Exactly-Once Semantics (EOS)

A complete guide to Kafka exactly-once semantics — delivery guarantees, idempotent producer, transactions, read_committed consumers, Kafka Streams EOS, zombie producer fencing, two-phase commit internals, and production patterns.

Kafka Log Compaction Explained

How Kafka log compaction preserves the latest value per key, enabling state stores, CDC changelog topics, and materialized views. Covers cleaner internals, tombstones, tiered storage, and KTable integration.

Kafka Performance Tuning Guide

End-to-end Kafka performance tuning — producer batching, broker I/O, consumer throughput, JVM and OS tuning, compression selection, tiered storage, and benchmarking methodology.

Kafka Producer

A **producer** is a client application that publishes (writes) messages to Kafka topics. It is responsible for:

Kafka Producers & Consumers

Combined guide to Kafka producers and consumers — internal architecture, delivery semantics, serialization, offset management, error handling, and production patterns.

Kafka Security Best Practices

End-to-end Kafka security guide covering authentication, ACLs, TLS 1.3 encryption, Zero Trust architecture, network isolation, credential management, and monitoring security events.

Kafka Streams — Complete Deep Dive

A comprehensive guide to Kafka Streams: from core concepts and internal architecture to stateful processing, failure recovery, exactly-once semantics, windowing, joins, interactive queries, and production system design patterns.

Kafka Topics

A topic is a named, durable stream of messages in Kafka — the logical category or feed where producers write and consumers read.

KRaft vs ZooKeeper: Kafka Metadata Architecture

Comprehensive guide comparing Apache Kafka's legacy ZooKeeper architecture with the modern KRaft (Kafka Raft) metadata mode — covering internal mechanics, failure scenarios, Strimzi Kubernetes deployment, migration strategies, and production deep dives for senior engineers.

Message Ordering with Partition Keys

Kafka guarantees **total ordering within a partition**. Messages written to the same partition are always consumed in the exact order they were produced.

Message Queues & Streaming

Guide to asynchronous messaging systems including Kafka, RabbitMQ, SQS, event sourcing, pub/sub patterns, consumer groups, ordering guarantees, and exactly-once semantics.

Parallel Consumer & Alternatives

Deep-dive into the Confluent Parallel Consumer model (now unmaintained), its internals, ordering modes, offset bitmap mechanics, and practical migration paths — Spring Boot virtual threads, manual executor dispatch, and the upcoming Kafka Share Groups (KIP-932).

Partitions

A partition is an ordered, immutable sequence of records within a topic — the fundamental unit of parallelism, replication, and storage scaling in Kafka.

Preventing Kafka Connect Rebalance Storms

Deep-dive into Kafka Connect rebalance storm internals — why they happen, the two rebalance protocols, incremental cooperative mechanics, task assignment algorithms, Spring Boot monitoring, and a production-grade rolling restart runbook.

Processing and Ordering

Kafka guarantees ordering within a partition, but single-threaded processing limits throughput. This guide covers four patterns for achieving high throughput while preserving per-key ordering.

Producer Acknowledgements (acks)

The `acks` configuration controls how many broker acknowledgements the producer requires before considering a send successful, trading off throughput, latency, and durability.

Raft Consensus Algorithm

A comprehensive guide to the Raft Consensus Algorithm — covering leader election, log replication, safety guarantees, and how it is implemented in Apache Kafka's KRaft metadata mode.

Real-Time Updates

Patterns for delivering real-time data to clients including WebSockets, Server-Sent Events, long polling, short polling, and push notification architectures.

Replication, ISR & Fault Tolerance

The replication factor defines how many copies of each partition exist across the cluster to guarantee fault tolerance and high availability.

Scaling Partitions

Partitions are the unit of parallelism in Kafka. Scaling them is critical for throughput but can break ordering for keyed topics. This guide covers the mechanics, risks, and migration strategies.

Scaling Writes

Deep-dive into high write throughput techniques — sharding, partitioning, WAL internals, LSM trees, async pipelines, batching, backpressure, idempotency, and distributed transactions — with production Java/Spring code and failure mode analysis.

Schema Registry

Deep-dive into Confluent Schema Registry — Avro wire format internals, subject naming strategies, compatibility mode mechanics, schema evolution patterns, code generation, topic reset with schema changes, and production Spring Boot configuration.

Transactional Outbox Pattern

A complete guide to the Transactional Outbox Pattern — from the Dual-Write problem for beginners to CDC vs polling internals, at-least-once guarantees, ordering semantics, and production monitoring for senior engineers.