[00:00] Why is Kafka fast? What is the secret? We'll talk about it in this video. Let's dive right in. We'll first start by acknowledging that the term fast is ambiguous. What does it even mean that Kafka is fast? Are we talking latency? Are we talking throughput? Is fast compared to what? [00:15] Kafka is optimized for high throughput. It is designed to move a large number of records in a short amount of time. Think of it as a very large pipe for moving liquid. The bigger the diameter of the pipe, the larger the volume of liquid that can move through it. So when someone [00:29] says Kafka is fast. They usually refer to Kafka's ability to move a lot of data efficiently. What are some of the design decisions that help Kafka move a lot of data quickly? There are many design decisions that contributed to Kafka's performance. In this video, we'll focus on two. We think these [00:44] two carry the most weight. The first one is Kafka's reliance on sequential I.O. Now what is sequential I.O.? Let's dig into it a little bit. There's a common misconception that this access is slow compared to memory access But this largely depends on the data access pattern There are two common disk access patterns random and sequential For hard drives it takes time to physically move the arm to different locations on the magnetic disks [01:08] This is what makes random access slow. For sequential access though, since the arm doesn't need to jump around, it is much faster to read and write blocks of data one after the other. Kafka takes advantage of this by using an append-only log as its primary data structure. An append-only [01:24] log adds new data to the end of the file. This access pattern is sequential. Now let's bring this idea home with some numbers. On modern hardware with an array of these hard disks, sequential writes can reach hundreds of megabytes per second, while random writes are measured in [01:39] hundreds of kilobytes per second. Sequential access is several orders of magnitude faster. Using hard disks has its cost advantage too. Compared to SSD, hard disks come as one-third of the price but with about three times the capacity giving Kafka a large pool of cheap disk [01:54] space without the performance penalty means that Kafka can cost effectively retain messages for a long period of time. This is a feature that was uncommon to messaging system before Kafka.