What the Wilcoxon signed-rank test is and when to use it
The Wilcoxon signed-rank test compares two related measurements on the same subjects, such as blood pressure before and after a change, or task time with an old and a new tool. It asks whether the typical paired difference is zero. Unlike the paired t-test, it does not assume the differences follow a bell curve. It only uses the size ordering of the differences and their signs, which makes it a good choice for small samples, skewed data, or measurements that contain a few extreme values.
Use it when each observation in one sample is naturally matched to exactly one observation in the other, and the differences can be meaningfully ranked from smallest to largest. If your two groups are independent (different people in each), the Mann-Whitney U test is the right tool instead. This calculator expects you to enter the differences directly, computed as after minus before for every pair, separated by commas or spaces.
How the calculation works
The method follows five steps, all of which the calculator performs on the differences you enter:
- Discard any difference equal to zero. The remaining count is n.
- Rank the absolute values of the differences from 1 (smallest) to n (largest). Tied absolute values each receive the average of the ranks they occupy.
- Add up the ranks that belong to positive differences (W+) and the ranks that belong to negative differences (W-). Their sum is always n(n+1)/2.
- Take the test statistic W = min(W+, W-).
- Under the null hypothesis, W has mean n(n+1)/4 and variance n(n+1)(2n+1)/24. The calculator forms z = (W - mean + 0.5) / SD, applying a 0.5 continuity correction toward the mean, and reports a two-tailed p-value from the normal distribution.
Because the p-value uses a normal approximation, it is most trustworthy when n is at least about 10 to 20. The calculator flags samples with fewer than six non-zero differences, where an exact critical-value table should be used instead. The approximation also does not apply the tie correction to the variance, so heavily tied data produces a slightly imprecise p-value.
Worked example
Suppose six participants each have an after-minus-before change of 2.1, -1.3, 0.8, 1.5, -0.4 and 3.2. None is zero, so n = 6. Sorted by absolute value the differences are 0.4 (negative), 0.8 (positive), 1.3 (negative), 1.5 (positive), 2.1 (positive), 3.2 (positive), receiving ranks 1 through 6 in that order.
The negative ranks are 1 and 3, so W- = 4. The positive ranks are 2, 4, 5 and 6, so W+ = 17 (check: 4 + 17 = 21 = 6 x 7 / 2). Therefore W = 4. The mean is 6 x 7 / 4 = 10.5 and the variance is 6 x 7 x 13 / 24 = 22.75, giving SD = 4.77. The deviation is 4 - 10.5 = -6.5; after the continuity correction it becomes -6.0, so z = -6.0 / 4.77 = -1.258 and the two-tailed p-value is 0.2084. The calculator prints exactly these values. Since 0.2084 exceeds 0.05, there is no evidence that the median difference is not zero, even though four of the six changes are positive.
Common mistakes and how to interpret the result
- Entering raw scores instead of differences. The tool needs one difference per pair. Pasting the before and after columns as a single list will produce meaningless ranks.
- Forgetting the sign convention. Reversing the subtraction swaps W+ and W-. W is the minimum of the two, so the p-value is unchanged, but your reading of the direction must match how you subtracted.
- Ignoring dropped zeros. Pairs with no change are excluded, so n can be smaller than the number of subjects. Many zeros signal low measurement resolution and reduce power.
- Treating a large p-value as proof of no effect. With small n the test has little power. A non-significant result means the data are inconclusive, not that the treatment did nothing.