---
title: 'The Design Choices Behind Kafka''s High Performance'
source: 'https://youtube.com/watch?v=la8tzEyg-hY'
video_id: 'la8tzEyg-hY'
date: 2026-09-03
duration_sec: 104
channel: 'ByteByteGo'
---

# The Design Choices Behind Kafka's High Performance

> Source: [The Design Choices Behind Kafka's High Performance](https://youtube.com/watch?v=la8tzEyg-hY)

## Summary

This video explains two fundamental design choices that give Apache Kafka its high performance: sequential I/O and the zero-copy principle. It contrasts the traditional data transfer path with Kafka's optimized approach, detailing how system calls and DMA reduce CPU involvement and copying overhead.

### Key Points

- **Zero copy principle** [00:16] — Modern Unix OS optimizes data transfer from disk to network, avoiding excess copying.
- **Without zero copy** [00:32] — Data goes: disk to OS cache → Kafka application → socket buffer → NIC buffer → network. Four copies and two system calls, inefficient.
- **With zero copy (sendfile)** [01:14] — Sends data directly from OS cache to network card buffer using sendfile system call, with only one copy.
- **DMA efficiency** [01:14] — Modern NICs use Direct Memory Access (DMA), copying without CPU involvement, boosting performance further.
- **Cornerstones** [01:30] — Sequential I/O and zero copy are the most important techniques, though Kafka uses others too.

## Transcript

The second design choice that gives Kafka its performance advantage is its focus on efficiency. Kafka moves a lot of data from network to disk and then from disk to network. It is critically important to eliminate excess copy when moving pages and pages of data between the disk and the
network. This is where zero copy principle comes into the picture. Modern Unix operating systems are highly optimized for transferring data from disk to network without copying data excessively. Let's dive deeper into see how this is done. First, we look at how Kafka sends a page of data on disk
to the consumer when zero copy is not used at all. First, the data is loaded from disk to the OS cache. Second, the data is copied from the OS cache into the Kafka application. Third, the data is copied
from Kafka to the socket buffer. And fourth, the data is copied from the socket buffer to the network interface card buffer. And finally, the data is sent over the network to the consumer. Now this is clearly inefficient. There are four copies and two system calls.
Now let's compare this to zero copy. The first step is the same. The data page is loaded from the disk to the OS cache. With zero copy, the Kafka application uses a system call called sendfile to tell the operating system to directly copy the data from the OS cache to the network interface card buffer.
In this optimized path, the only copy is from the OS cache into the network card buffer. With a modern network card, this copying is done with DMA. DMA is Direct Memory Access. When DMA is used, the CPU is not involved, making it even more
efficient. Now to recap, sequential I.O. and zero copy principle are the cornerstones to Kafka's high performance. Kafka uses other techniques to squeeze every ounce of performance out of the modern hardware, but these two are the most important in our view.
