About the Hamming Distance
The Hamming distance between two strings of equal length is the number of positions at which the corresponding characters differ. Named after Richard Hamming, who introduced it in his 1950 paper on error-detecting and error-correcting codes, it is a core tool in coding theory, cryptography, bioinformatics, and information theory whenever you need a simple count of "how many spots are different" between two same-length sequences.
The formula
For two sequences A and B, each of length n, the Hamming distance d(A, B) is defined as:
d(A, B) = number of positions i (1 to n) where A[i] ≠ B[i]
Two related figures follow directly from that count. The normalized Hamming distance divides the raw distance by the length, d(A, B) / n, giving a value from 0 (identical) to 1 (every position differs) that lets you compare sequences of different lengths on the same scale. The percent similarity is simply (1 − normalized distance) × 100, the share of positions that match.
A worked example
Compare the seven-character binary strings 1011101 and 1001001:
- Position 1: 1 vs 1 — match
- Position 2: 0 vs 0 — match
- Position 3: 1 vs 0 — differs
- Position 4: 1 vs 1 — match
- Position 5: 1 vs 0 — differs
- Position 6: 0 vs 0 — match
- Position 7: 1 vs 1 — match
Two of the seven positions differ, so the Hamming distance is 2, the normalized distance is 2/7 ≈ 0.2857, and the similarity is about 71.4%.
Why sequences must be the same length
Hamming distance compares characters position by position, so it is only defined when both sequences have identical length. If your two inputs have different lengths, there is no valid pairing — pad the shorter one (with zeros, spaces, or a placeholder character) or trim the longer one so both are equal length before comparing, or use a different metric such as edit (Levenshtein) distance, which is designed for strings of unequal length.
Where it is used
In telecommunications and storage, Hamming distance determines how many bit errors a code can detect or correct — a code with minimum distance d can detect up to d − 1 errors and correct up to ⌊(d − 1) / 2⌋ errors. In genetics, it counts point differences between equal-length DNA or protein sequences. In cryptography and information retrieval, it measures how similar two fixed-length hashes or fingerprints are.