---
title: 'Measuring the Entropy of English'
source: 'https://youtube.com/watch?v=-7etvZSBxlk'
video_id: '-7etvZSBxlk'
date: 2026-08-10
duration_sec: 145
---

# Measuring the Entropy of English

> Source: [Measuring the Entropy of English](https://youtube.com/watch?v=-7etvZSBxlk)

## Summary

This video explores Claude Shannon's pioneering work on measuring the entropy of English text, a foundational concept in information theory. It details his early statistical experiments, his innovative use of his wife Betty as a human predictor, and how these efforts led to estimates of the compressibility of English. The video connects these historical experiments to modern understanding of information limits.

### Key Points

- **Shannon's early experiments** [00:00] — Shannon began by tracking the statistics of letter sequences in English, such as which letters tend to follow 'TH', to understand the probability of each new letter.
- **Limitations of short sequences** [00:27] — Short sequences fail to capture context, missing rare letter combinations and reducing predictability, which is key for compression.
- **Using his wife as a predictor** [00:40] — To overcome data limitations, Shannon used his wife Betty to guess each next letter in a passage, recording her guesses to estimate information content.
- **Information as predictability** [01:06] — The idea was that a text string carries the same information if a perfect duplicate could fill it in, meaning information content depends on knowing the probability of each letter.
- **Refining the method** [01:20] — Shannon improved the approach by counting how many guesses were needed, rather than just whether they were correct, and combined this with statistical estimates.
- **Entropy estimate** [01:46] — With at least 100 characters of context, English text can be compressed to around 1 bit per character, a key result from Shannon's experiments.
- **Beyond pure data** [01:59] — Shannon was forced to go beyond pure data to estimate entropy, highlighting the need for human-based prediction in information theory.
- **Modern relevance** [02:14] — Today, more than 75 years later, approaching this limit still involves probing human prediction, not just statistical analysis.

### Conclusion

Shannon's work on measuring the entropy of English, using both statistical analysis and human prediction, established a fundamental limit on text compressibility. This historical experiment remains relevant to modern information theory and compression algorithms.

## Transcript

compressibility of English text depends on how at understanding the probability of each new His earliest experiments involved looking at
tracking the statistics of what tended to follow. where you see the letters TH and build The problem here is that this completely
most notably those that never show up But longer sequences give more context to guide its most predictable and hence most compressible.
of language, he needed some other way So he turned to one of the most readily to him in the 1940s: His wife, Betty.
each new letter from a given passage. and every time she guessed correctly, His idea was that this new string of text
the same information, in the sense that it duplicate of his wife to fill in the entire text. of information requires knowing the
so he needed something better. of Printed English, Shannon outlined an instead of just logging whether their guess
many guesses were necessary for his human Separately, he combined the idea of statistics to make an estimate at the implicit probabilities
letter, based on their number of guesses. least 100 characters of context, be compressible down to around 1 bit per
estimate, he was forced to go beyond pure data Today, more than 75 years later, to this limit is not through merely probing at
This all comes from the first video If you want more, take a look at the channel.
