Correlation and Regression

Stage 16 of 23 Strand 5 of 5 8 lessons

8 illustrated lessons, each teaching the why before the how.

Revise Correlation and Regression with flashcards →

Jump to a lesson

Pearson’s Correlation Coefficient

One number from minus one to one.

Pearson’s r measures how closely two measurements follow one straight line, from minus one to one

Every dot on one rising line gives r = 1, the largest value r can take.

Every dot on one falling line gives r = −1. The sign of r is the direction.

A shapeless cloud gives r near 0: one reading says nothing about the other.

The same rise with more scatter drops r toward 0. The size of r is the strength.

So one number carries both: the sign gives the direction, the size gives the strength.

And r judges straight lines only — this tight arch still gives r near 0.

Now you

Which value of r fits these dots best?

Critical Values of r

Small samples look linear by luck alone.

A correlation counts as evidence only once r beats the critical value for that sample size

Three points nearly always look linear, so a high r on a tiny sample means little.

A table gives the value r must beat for that sample size at the 5 percent level.

Ten pairs put the bar at 0.632, and an r of 0.81 clears it comfortably.

The null claim has a name: H₀ says ρ, the population correlation, is zero.

Falling short settles nothing — ten pairs simply cannot rule chance out.

The bar assumes both readings are normal: an even cloud, no bend, and no outliers.

Now you

With n = 20 the critical value is 0.444, and r comes to 0.564. What follows?

With n = 10 the critical value is 0.632, and r comes to 0.512. What follows?

The Least-Squares Regression Line

The line that makes the squared gaps smallest.

The regression line of y on x is the line that makes the squared vertical gaps as small as possible

Five dots rise from left to right. Many straight lines pass near them, and one fits best.

Measure the vertical gap from each dot to the line. Those gaps are the residuals.

Tilt the line and some gaps shrink while others grow. Gaps above and below the line have opposite signs, so square them before adding, or they cancel.

One line makes the sum of the squared residuals smallest, and that one is the fit.

Write the line as y = ax + b. A calculator gives a and b from the list of pairs.

Read them in context: a is the rise per unit of x, and b is y when x is 0.

Now you

Which total does the regression line of y on x make smallest?

Cost = 8n + 35 was fitted to n items. What is the 35?

Predicting from a Regression Line

Inside the data a result, beyond it a guess.

A regression prediction is a result inside the range of the data and a guess beyond it

The data run from 2 hours to 8. The line was fitted to that range alone.

Read the line at 6 hours and measured points sit on either side. That is interpolation.

Stretch the same line to 18 hours, and no data point supports it there.

So reading inside the range is interpolation; beyond it is extrapolation.

The line of y on x predicts y from x only, because its gaps were measured in y.

Now you

The fitted line is y = 0.5x + 1.5. Predict y when x is 3.

The fitted line is y = 0.5x + 3. Predict y when x is 4.

Spearman’s Rank Correlation

Rank the values first, then correlate the ranks.

Spearman’s rank correlation replaces every value by its rank and measures how well the two orders agree

These points rise at every step, yet they do not follow a straight line.

Replace each value by its rank, its place in size order: 1 for the smallest, then 2, 3, 4.

Rank both measurements, then correlate the two rank columns instead of the values.

Ranking turns the curved pattern into a straight line, and two orders that agree perfectly give a rank correlation of 1.

Equal values take the average of the places they cover, so a tie for 2nd and 3rd gets rank 2.5 each.

One outlier drags r a long way, but it can only move a rank by a place or two.

Use r for a straight-line pattern, and rank correlation when only the order matters.

Now you

Two values tie for the places 2 and 3. What rank does each take?

Ranked smallest first, what rank does the marked value get?

The Regression Line of x on y

The other fit, for when x is the unknown.

The line of x on y makes the horizontal gaps smallest, and it is the line that predicts x

The line of y on x minimizes the vertical residuals, so it predicts a score.

Swap the axes and fit again: the residuals are now measured horizontally.

Back on the original axes, the fit of x on y is a different line.

Write x on y as x = cy + d. Put a score into it and it gives back the hours.

Both lines cross at the mean point, and they coincide only when r is 1 or −1.

Choose by what you are predicting: minimize the residuals in that variable.

Now you

Where do the line of y on x and the line of x on y always meet?

You know a score and want the hours behind it. Which line?

The Coefficient of Determination

The share of the variation a model accounts for.

R squared is the share of the variation in y that the fitted model accounts for

With no model at all, the best guess for y is its mean. Those squared gaps total 18.

The regression line leaves smaller residuals. Squared and added, they total 3.6.

R squared is one minus the unexplained share, so 0.8 of the variation is accounted for.

For a straight line R squared is exactly r squared, so it never carries a sign.

Adding terms can only raise R squared, so the largest value is not always the best fit.

here is 0.91, but the residuals still curve, so look at them as well as at .

Now you

The squared gaps to the mean total 25 and the squared gaps to the line total 4. What is R squared?

A model reports R² = 0.81. What does that say?

Non-Linear Regression

The same least-squares rule, over a curve.

Least squares fits a curve by the same rule that fits a line, and the shape of the data decides the family

The best straight line here sits below the middle dots and above the outer ones.

Fit a quadratic by the same rule and every gap closes: a straight line was the wrong shape for these data.

A calculator offers five families: quadratic, cubic, exponential, power and sine.

Counts that double every two weeks need ab^x, a curve that never turns back down.

The rule never changes. What changes is the family of curves least squares searches over.

Now you

Which quantity does least squares make smallest when it fits a curve?

Counts that double every two weeks are fitted best by which family?

Continue your journey in the app — save your progress