Coefficient of Determination Formula

Coefficient of determination is the proportion of the total variation in the response variable y that is explained by the linear relationship with the explanatory variable x.

The Formula

r2=1−SSresidualSStotal=1−∑(yi−y^i)2∑(yi−yˉ)2

When to use: Total variation in y has two parts: what the regression line explains and what's left over (residual variation). If r2=0.85, the regression line accounts for 85% of why y values differ from each other, and 15% is unexplained. Think of r2 as a report card for how well x predicts y.

Quick Example

Correlation between study hours and test score: r=0.9. r2=0.81 Interpretation: 81% of the variation in test scores is explained by the linear relationship with study hours.

Notation

r2 ranges from 0 to 1. SStotal = total sum of squares. SSresidual = residual sum of squares.

What This Formula Means

The proportion of the total variation in the response variable y that is explained by the linear relationship with the explanatory variable x. It equals the square of the correlation coefficient: r2.

Total variation in y has two parts: what the regression line explains and what's left over (residual variation). If r2=0.85, the regression line accounts for 85% of why y values differ from each other, and 15% is unexplained. Think of r2 as a report card for how well x predicts y.

Formal View

r2=1−SSresSStot=1−∑(yi−y^i)2∑(yi−yˉ)2 where 0≤r2≤1

Worked Examples

Example 1

medium
A regression model has SST=500 (total variation) and SSE=125 (unexplained variation). Calculate R2 and interpret its meaning.

Answer

R2=0.75. The model explains 75% of variation in y.

First step

1
R2=1−SSESST=1−125500=1−0.25=0.75

See the full worked solution + why-it-works coaching

SetupKey insightWhy it worksCommon pitfallConnection

Unlock answer keys One Family plan — every worked solution, all subjects

Example 2

hard
Two models predict house prices: Model 1 (size only): R2=0.60. Model 2 (size + neighborhood + age): R2=0.85. Explain what the increase in R2 means and what caution should be applied with multi-variable R2.

Example 3

medium
A model has SST=1200 and r2=0.7. Find the explained sum of squares SSR and the residual SSE.

Common Mistakes

  • Reporting r when the question asks for r2 - square the correlation; r=0.7 gives r2=0.49, not 0.7.
  • Reading r2 as causation - it measures explained variation, never that x causes y.
  • Letting r2 go negative or above 1 - it's a proportion between 0 and 1, so any value outside that range is an error.

Why This Formula Matters

r2 is the standard one-number report card for a regression's predictive usefulness, and squaring r exposes how much weaker a 'decent' correlation really is (r=0.7 explains only 49%). Mixing it up with r or with causation is what leads people to overstate how much a model actually tells them. Recognizing it by "Am I reporting the fraction of y's variation explained by the linear model (a 0-to-1 number), not the slope or the correlation's sign?" — rather than by familiar numbers — is what lets a student tell it apart from correlation r and slope b and residual variation in a mixed problem set.

Frequently Asked Questions

What is the Coefficient of Determination formula?

The proportion of the total variation in the response variable y that is explained by the linear relationship with the explanatory variable x. It equals the square of the correlation coefficient: r2.

How do you use the Coefficient of Determination formula?

Total variation in y has two parts: what the regression line explains and what's left over (residual variation). If r2=0.85, the regression line accounts for 85% of why y values differ from each other, and 15% is unexplained. Think of r2 as a report card for how well x predicts y.

What do the symbols mean in the Coefficient of Determination formula?

r2 ranges from 0 to 1. SStotal = total sum of squares. SSresidual = residual sum of squares.

Why is the Coefficient of Determination formula important in Math?

r2 is the standard one-number report card for a regression's predictive usefulness, and squaring r exposes how much weaker a 'decent' correlation really is (r=0.7 explains only 49%). Mixing it up with r or with causation is what leads people to overstate how much a model actually tells them. Recognizing it by "Am I reporting the fraction of y's variation explained by the linear model (a 0-to-1 number), not the slope or the correlation's sign?" — rather than by familiar numbers — is what lets a student tell it apart from correlation r and slope b and residual variation in a mixed problem set.

What do students get wrong about Coefficient of Determination?

The procedure for coefficient of determination is the easy part; the trap is reporting r when the question asks for r2. Asking "Am I reporting the fraction of y's variation explained by the linear model (a 0-to-1 number), not the slope or the correlation's sign?" first is what keeps a correct-looking calculation from being attached to the wrong concept.

What should I learn before the Coefficient of Determination formula?

Before studying the Coefficient of Determination formula, you should understand: correlation, linear regression lsrl, residuals.