/ THE IDEA
An outcome carries more information when it was less expected under the chosen model. Drawing a common red ball from a bag tells you little; drawing its only blue ball tells you more. Entropy averages that surprise over possible messages. A fair coin is maximally unpredictable: heads and tails each need one bit on average. A coin that is heads 99% of the time can be encoded more efficiently across a long sequence because tails is rare.
THE FORMAL IDEA
H = −Σ p(x) log₂ p(x)
| p(x) = probability of message x | | −log₂p(x) = surprise measured in bits | | Σ averages that surprise across all possible messages |
|
RUN THE TINY EXAMPLE
Change a coin’s predictability
Fair coin: p(heads)=0.5 → H = 1 bit per toss 90% heads: H = −[0.9 log₂(0.9) + 0.1 log₂(0.1)] ≈ 0.47 bits per toss More predictable source → lower theoretical average description length
|
You cannot literally send half a bit for one toss, but over thousands of tosses a good code can pack common sequences efficiently and approach the average.
/ SO WHAT?
This explains why compression finds repeated language and image patterns, while good encryption removes visible predictability. In statistical physics, a closely related entropy measures how many hidden microscopic arrangements fit the same visible state.
ONE CAVEAT |
| Entropy depends on the probability model. A compressor that has not learned the pattern may miss it, and real files carry headers, error protection and finite-size overhead above the theoretical limit. |
KEEP THIS
The harder a source is to predict, the more bits its messages need on average.
|
NEXT: How one program juggles many unfinished jobs
|