Apache Kafka Tutorial 0/111 lessons ~6 min read Lesson 53
Retention Policies
What is Retention Policies?
Course progress0%
Focus
9 guided sections
Practice signal
Examples included
Career prep
Foundation builder
Introduction
What is Retention Policies? Retention controls how long Kafka keeps records by time, size, or compaction.
Understanding the topic
What happens — Retention Policies:
- Plan partitions and replication factor.
- Configure retention and compaction.
- Monitor ISR, lag, disk.
- Scale brokers and partitions with load.
| Term | Description |
|---|---|
| Consumer group | A set of consumers sharing work — each partition goes to at most one member at a time. |
| Offset | Position in a partition log — where this consumer last read. |
| Poll loop | consumer.poll() fetches batches of records; keep processing faster than max.poll.interval.ms. |
| Commit | Saving offset to Kafka after processing — sync or async, manual or auto. |
| Lag | Difference between latest offset and consumer offset — key health metric. |
Visual explanation
Pipeline view:
text
Consumer polls records from topic partitions↓Process business logic (DB, API, etc.)↓Commit offset after success↓Monitor consumer lag
Step-by-step explanation
- Subscribe — Consumer joins a group and receives partition assignments.
- Poll — Fetch records in batches with consumer.poll().
- Process — Run business logic for each record.
- Commit — Save offset after successful side effects.
- Repeat — Continue polling; rebalance if group membership changes.
Informative example
Example:
bash
kafka-topics --bootstrap-server localhost:9092 --create --topic retention-policies --partitions 6 --replication-factor 3kafka-console-producer --bootstrap-server localhost:9092 --topic retention-policieskafka-console-consumer --bootstrap-server localhost:9092 --topic retention-policies --from-beginningkafka-consumer-groups --bootstrap-server localhost:9092 --describe --group retention-policies-service
Execution workflow
1Retention Policies workflow
1 / 5Subscribe
Consumer joins a group and receives partition assignments.
Best practices
- Use manual commits for critical side effects.
- Make consumers idempotent with event IDs or unique constraints.
- Keep processing under max.poll.interval.ms or use pause/resume patterns.
- Send poison messages to DLT with failure metadata.
Common mistakes
- Committing before database writes complete.
- Scaling consumers beyond partition count and expecting more throughput.
- Blocking the poll loop with slow downstream calls.
Summary
Retention Policies — Commit offsets only after successful processing; make handlers idempotent; monitor lag per partition.
Ready to mark this lesson complete?Track your journey across the entire course.