Hypergeometric Distribution Calculator - Sampling Without Replacement

Calculate hypergeometric probabilities for sampling without replacement from finite populations.

Results

Calculated
Result
—

What the hypergeometric distribution is for

The hypergeometric distribution answers a specific kind of probability question: if you draw a fixed-size sample from a finite population without replacement, what's the chance of getting exactly a certain number of "successes" (items with the trait you care about)? Classic examples include drawing cards from a deck, pulling defective parts from a batch, selecting winning tickets from a fixed pool, or auditing a sample of records for errors. The defining feature is that each draw changes the composition of what's left — once you remove a red marble from the bag, there's one fewer red marble (and one fewer marble overall) for the next draw.

This is what separates it from the binomial distribution, which assumes each trial is independent with a constant success probability (sampling with replacement, or drawing from an effectively infinite population). Use the hypergeometric distribution whenever your population is small enough, relative to your sample, that removing items measurably changes the odds for what's left — quality control on a finite production lot, dealing cards, or picking a committee from a small group are typical cases.

The formula and what each variable means

P(X = k) = [C(K, k) × C(N−K, n−k)] ÷ C(N, n)

Where N is the population size (total items), K is the number of "success" items in the whole population, n is the sample size you draw, k is the number of successes you want the probability for, and C(a, b) is the combination "a choose b" — the number of ways to choose b items from a items, ignoring order. The numerator counts the ways to choose exactly k successes from the K available successes and the remaining n−k draws from the N−K non-successes; the denominator counts every possible way to choose any n items from the full population N. The calculator also reports the cumulative probability P(X ≤ k), the mean (μ = nK/N), and the variance and standard deviation, which describe the full shape of the distribution around that mean.

Worked example

Suppose a bag has 20 marbles total, 6 of them red, and you draw 5 marbles without putting any back. What's the probability of drawing exactly 2 red marbles? Here N=20, K=6, n=5, k=2. The count of ways to pick 2 reds from 6 is C(6,2)=15. The count of ways to pick the remaining 3 draws from the 14 non-red marbles is C(14,3)=364. The count of ways to pick any 5 marbles from 20 is C(20,5)=15,504. So P(X=2) = (15 × 364) ÷ 15,504 = 5,460 ÷ 15,504 ≈ 35.22%. Entering N=20, K=6, n=5, k=2 into the calculator above reproduces this exact-probability figure, along with a cumulative P(X≤2) of roughly 86.9%, a mean of 1.5000 (5 × 6 ÷ 20), and a standard deviation around 0.91.

Common mistakes and how to interpret the result

  • Using the binomial distribution when sampling without replacement from a small population. If your sample is more than roughly 5-10% of the population, treating draws as independent (binomial) noticeably overstates or understates probabilities compared to the correct hypergeometric result.
  • Confusing "exact" and "cumulative" probability. P(X = k) is the chance of exactly k successes; P(X ≤ k) sums that probability with every smaller outcome. Use cumulative probability for "at most" or "at least" questions, not the exact PMF value alone.
  • Entering K or n larger than N. The number of successes in the population can't exceed the population itself, and you can't sample more items than exist — the calculator flags these cases rather than returning a nonsensical probability.
  • Forgetting the population shrinks as the sample grows. As n approaches N, remaining uncertainty drops sharply — sampling nearly the whole population makes almost any specific outcome close to certain or impossible, which is exactly the "without replacement" effect this distribution captures.

Frequently Asked Questions

How is the hypergeometric distribution different from the binomial distribution?
The binomial distribution assumes each draw is independent with a fixed success probability, which is accurate when sampling with replacement or from a very large population. The hypergeometric distribution accounts for sampling without replacement from a finite population, where each draw changes the odds for the next one — the smaller the population relative to the sample, the bigger the difference between the two.
What does C(N, n) mean in the formula?
C(N, n), read as "N choose n," is the number of ways to select n items from a group of N items where order doesn't matter. It's calculated as N! ÷ (n! × (N−n)!), and this calculator computes it internally using logarithms of factorials to stay numerically stable for large populations.
When should I use the cumulative probability instead of the exact one?
Use the exact probability P(X = k) when you want the chance of a specific count. Use the cumulative probability P(X ≤ k) when the question is phrased as "at most k" successes, or when you need P(X ≥ k) — which you get by computing 1 minus the cumulative probability up through k−1.
What do the mean and variance tell me?
The mean (nK/N) is the expected number of successes you'd see on average if you repeated the sample many times. The variance and standard deviation describe how much that count typically varies from draw to draw — a small variance means outcomes cluster tightly around the mean, while a larger one means more spread.