What linear regression does and when to use it
Linear regression finds the straight line that best fits a set of paired (x, y) data points — the line that minimizes the total squared vertical distance between itself and every point. It's one of the most widely used tools in statistics and data analysis because it answers two practical questions at once: is there a linear relationship between two variables, and if so, what does that relationship predict for an x value you haven't observed? This calculator handles both — it fits the line and, if you supply a value, predicts the corresponding y.
Use it when you have paired numeric measurements and want to quantify a trend: study hours versus test scores, advertising spend versus sales, temperature versus energy usage, or any two variables you suspect move together. It's built for simple linear regression (one predictor variable) using the ordinary least squares method taught in introductory statistics — it isn't meant for data with a clearly curved (non-linear) relationship, which needs a different model entirely.
The formula and what each term means
Best fit line: y = mx + b, where the slope and intercept are found from:
m (slope) = [n·Σ(xy) − Σx·Σy] ÷ [n·Σ(x²) − (Σx)²]
b (intercept) = [Σy − m·Σx] ÷ n
Here n is the number of paired data points, Σx and Σy are the sums of all x and y values, Σ(xy) is the sum of each x multiplied by its paired y, and Σ(x²) is the sum of each x value squared. Once the line is fit, R² (the coefficient of determination) measures how well it fits: R² = 1 − (sum of squared residuals ÷ total sum of squares around the mean of y). R² ranges from 0 (the line explains none of the variation in y) to 1 (the line passes through every point exactly).
Worked example
Using the default data — X: 1, 2, 3, 4, 5 and Y: 2.1, 3.9, 6.2, 7.8, 10.1 — n=5, Σx=15, Σy=30.1, Σ(xy)=110.2, Σ(x²)=55. Slope: m = (5×110.2 − 15×30.1) ÷ (5×55 − 15²) = (551 − 451.5) ÷ (275 − 225) = 99.5 ÷ 50 = 1.99. Intercept: b = (30.1 − 1.99×15) ÷ 5 = (30.1 − 29.85) ÷ 5 = 0.05. The best-fit line is y = 1.99x + 0.05, with R² ≈ 0.9973 (a very strong linear fit). Predicting y at x=6: y = 1.99×6 + 0.05 = 11.99. Entering these exact X and Y values with a predict-X of 6 reproduces every one of these numbers in the calculator above.
Common mistakes and how to interpret the result
- Trusting a high R² without checking the data visually. R² measures how closely points follow a straight line, but a clearly curved relationship can still occasionally show a moderate R² — plot your data or reason about it before assuming linearity is the right model.
- Extrapolating far beyond your data range. A line fit on x-values from 1 to 5 is only well-supported in that range; predicting y at x=500 assumes the same relationship holds far outside where you have evidence, which is often false.
- Confusing correlation with causation. A strong linear fit between two variables shows they move together, not that one causes the other — a third factor could be driving both.
- Letting a single outlier dominate the line. Least squares regression is sensitive to outliers because it squares each residual — one far-off point can pull the slope and intercept noticeably away from what the rest of the data suggests.