Scatter Plot Calculator

Enter your paired (x, y) data points to fit a least-squares best-fit line and get the slope, intercept, and Pearson correlation coefficient describing the trend.

Quick Facts

Best-fit line
y = mx + b
Least-squares slope m = (nΣxy − ΣxΣy) / (nΣx² − (Σx)²); intercept b = (Σy − mΣx) / n.
Correlation coefficient
r = (nΣxy − ΣxΣy) / √[(nΣx² − (Σx)²)(nΣy² − (Σy)²)]
Ranges from −1 (perfect negative trend) to 1 (perfect positive trend); 0 means no linear trend.
Coefficient of determination
r²
The share of the variation in y explained by the line, expressed as a percentage.

Your Results

Calculated
Slope (m)
-
Change in y per unit x
Y-Intercept (b)
-
Predicted y when x = 0
Correlation (r)
-
Strength and direction of trend
R² (Determination)
-
% of variance explained by the line

Ready

Enter your x, y data points (one pair per line), then press Calculate.

How the Scatter Plot Calculator works

A scatter plot shows every (x, y) data point as a dot, and the pattern those dots make reveals whether two variables move together. This calculator turns that visual pattern into numbers: it fits the straight line that best matches your points using the least-squares method, then measures how tightly the points cluster around that line using the Pearson correlation coefficient. Together, the best-fit line and the correlation coefficient are the standard numeric summary of a scatter plot's trend.

Formula and method

For n data points, the least-squares regression line is y = mx + b, where the slope and intercept are computed directly from the sums of the data:

m = (nΣxy − ΣxΣy) / (nΣx² − (Σx)²)   and   b = (Σy − mΣx) / n

This line minimizes the sum of the squared vertical distances between the points and the line — no other straight line has a smaller total squared error. The Pearson correlation coefficient r then measures how closely the points hug that line:

r = (nΣxy − ΣxΣy) / √[(nΣx² − (Σx)²)(nΣy² − (Σy)²)]

r always falls between −1 and 1. A value close to 1 means a strong positive (upward) trend, close to −1 means a strong negative (downward) trend, and close to 0 means the points are scattered with no clear linear pattern. Squaring r gives r² (the coefficient of determination), which reports the percentage of the variation in y that is explained by x through the line — for example, r = 0.9 gives r² = 0.81, meaning the line explains 81% of the spread in y.

Common sources of error

  • Too few points: a line always fits exactly through 2 points, so 2-point "trends" are meaningless — use enough points to see a genuine pattern.
  • Reading correlation as causation: a strong r shows that x and y move together, not that x causes y — a third factor could drive both.
  • Ignoring curvature or outliers: the least-squares line only captures a straight-line relationship; a single extreme outlier or an obviously curved pattern can produce a misleading slope and r.
  • Vertical spread of x: if every x-value is identical, there is no unique slope (a vertical "line" has undefined slope) — the calculator will flag this case.

Interpreting your result

Read the slope as "y changes by m for every 1-unit increase in x," and the intercept as the predicted y-value when x = 0 (which may or may not be meaningful depending on your data). Use |r| as a rough strength guide: 0.7–1.0 is a strong linear relationship, 0.3–0.7 is moderate, and below 0.3 is weak. Always look at the actual scatter of points too — r only measures straight-line association, so it can be small even when a strong curved relationship exists.

Applications

Scatter plot and regression analysis are used to check whether study hours predict test scores, whether advertising spend predicts sales, whether temperature predicts ice cream sales, and in general to quantify the strength of any relationship between two measured quantities before deciding whether to trust or act on that relationship.

Frequently Asked Questions

What does a scatter plot calculator actually compute?
It takes your list of (x, y) data points and fits the least-squares regression line y = mx + b through them, then reports the slope, the y-intercept, and the Pearson correlation coefficient r, which together summarize the linear trend you would see by eye on a scatter plot.
How do I read the correlation coefficient r?
r ranges from −1 to 1. Values near 1 mean a strong positive relationship (y rises as x rises), values near −1 mean a strong negative relationship (y falls as x rises), and values near 0 mean the points show little to no linear pattern.
What is the difference between r and r²?
r measures the direction and strength of the linear relationship (−1 to 1). r², the coefficient of determination, is r multiplied by itself and reports the share of the variation in y (as a percentage) that is explained by the fitted line, regardless of direction.
How many data points do I need for a meaningful result?
The math only requires 2 points to draw a line, but with just 2 points the fit is exact and meaningless as a trend. Use at least 5–6 points, and treat r and the slope as more reliable the more points you include and the more spread out the x-values are.