What an F-Test Is and When to Use It
An F-test for two variances checks whether two samples come from populations with equal variance (spread), rather than equal mean. It's the test you reach for before choosing between a "Student's" two-sample t-test (which assumes equal variances) and Welch's t-test (which doesn't) — many textbooks recommend running an F-test first to decide which t-test formula applies. It's also the building block behind ANOVA, which extends the same variance-ratio logic to compare more than two groups at once.
The test works by forming a ratio of the two sample variances. If the two populations really do have equal variance, that ratio should cluster around 1; values far from 1 (in either direction) suggest the variances genuinely differ. Typical uses include comparing measurement consistency between two instruments or two lab technicians, checking whether a manufacturing process's variability changed after an intervention, or verifying the equal-variance assumption before a t-test.
The Formula
The F-statistic is the ratio of the two sample variances, with the larger variance placed on top for a two-tailed test so that F is always ≥ 1:
F = s1² ÷ s2² (larger variance ÷ smaller variance, for a two-tailed test)
- s1², s2² — the sample variances of group 1 and group 2
- n1, n2 — the sample sizes of group 1 and group 2
- df1 = n1 − 1, df2 = n2 − 1 — the degrees of freedom for the numerator and denominator
F is then compared against a critical value from the F-distribution with (df1, df2) degrees of freedom at your chosen significance level (alpha). For a two-tailed test, the critical value uses alpha/2 in each tail because either sample's variance could be the larger one; for the one-tailed test (testing specifically whether variance 1 exceeds variance 2), the full alpha is used in a single tail.
Worked Example
Using the calculator's own default inputs — Sample 1: variance 25.4, n1 = 15; Sample 2: variance 14.8, n2 = 12; alpha = 0.05, two-tailed:
- df1 = 15 − 1 = 14, df2 = 12 − 1 = 11
- Since 25.4 > 14.8, F = 25.4 ÷ 14.8 = 1.7162
- Critical F-value at alpha = 0.05 (two-tailed, df = 14, 11) = 3.3588
- P-value = 0.3729
Because F (1.7162) is less than the critical value (3.3588), and the p-value (0.3729) is well above 0.05, the calculator correctly reports "Not Statistically Significant" — there isn't enough evidence from these two samples to conclude their population variances differ.
Common Mistakes / How to Interpret the Result
- Using the one-tailed option when you don't have a directional hypothesis in advance. The one-tailed test should only be used when you specifically predicted, before seeing the data, that sample 1's variance would be larger — using it after the fact to get a smaller p-value is a form of p-hacking.
- Confusing "not significant" with "the variances are equal." Failing to reject the null hypothesis means the data didn't provide strong enough evidence of a difference — it does not prove the variances are truly identical, especially with small samples.
- Ignoring the normality assumption. The F-test for variances is quite sensitive to non-normal data (unlike the t-test for means, which is fairly robust). If your data is heavily skewed or has outliers, consider Levene's test instead, which is more robust to non-normality.
- Mixing up degrees of freedom order. df1 (numerator) always corresponds to whichever sample's variance is on top of the ratio — swapping df1 and df2 when reading a critical-value table will give the wrong critical value, since the F-distribution is not symmetric between its two degrees-of-freedom parameters.