Apache Kafka Tutorial 0/111 lessons ~6 min read Lesson 80

    Distributed Transactions

    What is Distributed Transactions?

    Course progress0%
    Focus
    9 guided sections
    Practice signal
    Examples included
    Career prep
    Foundation builder

    Introduction

    What is Distributed Transactions? Distributed transactions across services are usually replaced by sagas, outbox, idempotency, and compensating events.

    Understanding the topic

    What happens — Distributed Transactions:

    • Begin transaction with transactional.id.
    • Send to multiple partitions atomically.
    • Commit or abort as one unit.
    • Used for exactly-once stream processing.
    TermDescription
    TopicA named channel where producers publish and consumers read — e.g. user_signups or payment_logs.
    PartitionA topic is split into partitions for parallel storage and processing. Ordering is per partition.
    ProducerRecordThe object you send: topic name, optional key, value (payload), and optional headers.
    SerializationConverting your data (String, JSON, Avro) into bytes — Kafka only stores bytes.
    KeyOptional field used to pick a partition. Same key → same partition → ordered delivery for that key.
    BrokerKafka server that stores messages and serves producers and consumers.
    Ack (Acknowledgment)How many replicas must confirm receipt before the send is considered successful (acks=0, 1, or all).

    Visual explanation

    Pipeline view:

    text
    App creates message (e.g. "User signed up")
    Producer sends to topic (e.g. user_signups)
    Kafka stores in partitioned log (replicated)
    Consumers / analytics / alerts read in real time

    Step-by-step explanation

    1. Initialization — Producer connects to bootstrap servers and receives metadata (topics, partitions, leaders).
    2. Message creation — Application builds a ProducerRecord with topic, key, value, and optional timestamp.
    3. Serialization — Key and value serializers convert objects to bytes (e.g. StringSerializer).
    4. Partitioning — Hash of key picks partition; no key → round-robin / sticky partitioner.
    5. Batching — Records accumulate in the RecordAccumulator per partition for efficient network use.
    6. Sending — Background sender thread ships batches to the partition leader broker.
    7. Acknowledgement — Broker responds; producer retries on failure based on acks and retry settings.

    Informative example

    Example:

    bash
    kafka-topics --bootstrap-server localhost:9092 --create --topic distributed-transactions --partitions 6 --replication-factor 3
    kafka-console-producer --bootstrap-server localhost:9092 --topic distributed-transactions
    kafka-console-consumer --bootstrap-server localhost:9092 --topic distributed-transactions --from-beginning
    kafka-consumer-groups --bootstrap-server localhost:9092 --describe --group distributed-transactions-service

    Execution workflow

    1Distributed Transactions workflow
    1 / 7

    Initialization

    Producer connects to bootstrap servers and receives metadata (topics, partitions, leaders).

    Best practices

    • Use acks=all with min.insync.replicas for critical events.
    • Enable idempotence for retry safety.
    • Choose keys intentionally; random keys destroy ordering.
    • Monitor record-error-rate, request-latency, batch-size, and retries.

    Common mistakes

    • Ignoring asynchronous send failures.
    • Using null keys for events that require aggregate ordering.
    • Making producer timeouts too low and failing during normal broker leader movement.

    Summary

    Distributed Transactions — Use acks=all, enable idempotence, choose meaningful keys, and handle send callbacks for errors.

    Ready to mark this lesson complete?Track your journey across the entire course.