Inference for Regression Formula

Inference for regression is using hypothesis tests and confidence intervals to draw conclusions about the true population slope β_1 of the linear relationship y = β_0 + β_1 x + ε, based on sample data.

The Formula

t=b−β1,0SEbwhereSEb=s∑(xi−xˉ)2

When to use: You computed a sample regression line with slope b=2.3. But is the true population slope actually different from zero? Maybe there's really no linear relationship and you just got a slope by chance. The regression t-test asks: 'Is my sample slope far enough from zero that it's unlikely to have occurred by random variation alone?'

Quick Example

Sample slope b=2.3, SEb=0.8, n=25. t=b−0SEb=2.30.8=2.875(df=23) The p-value ≈0.008<0.05, so reject H0:β1=0. There is evidence of a linear relationship.

Notation

b = sample slope, β1 = population slope, SEb = standard error of the slope, s = standard deviation of residuals, df=n−2.

What This Formula Means

Using hypothesis tests and confidence intervals to draw conclusions about the true population slope β1 of the linear relationship y=β0+β1x+ε, based on sample data.

You computed a sample regression line with slope b=2.3. But is the true population slope actually different from zero? Maybe there's really no linear relationship and you just got a slope by chance. The regression t-test asks: 'Is my sample slope far enough from zero that it's unlikely to have occurred by random variation alone?'

Formal View

t=b−β1,0SEb with df=n−2 where SEb=s∑(xi−xˉ)2; CI: b±t∗⋅SEb

Worked Examples

Example 1

medium
A regression output shows: slope b=2.5, SEb=0.8, n=30. Test H0:β=0 vs Ha:β≠0 at α=0.05 using a t-test.

Answer

t=3.125>2.048. Reject H0. The slope is statistically significant at α=0.05.

First step

1
Test statistic: t=b−β0SEb=2.5−00.8=3.125

See the full worked solution + why-it-works coaching

SetupKey insightWhy it worksCommon pitfallConnection

Unlock answer keys One Family plan — every worked solution, all subjects

Example 2

hard
Construct a 95% confidence interval for the slope β given: b=1.8, SEb=0.5, n=25, and t0.025,23∗=2.069.

Example 3

medium
A study gives b=3.2, SEb=1.0, n=12. With t0.025,10∗=2.228, construct the 95% CI for the slope.

Common Mistakes

  • Treating a nonzero sample slope as proof of a population relationship - test b against its standard error before concluding β1≠0.
  • Using the wrong degrees of freedom - regression inference uses df=n−2, not n−1.
  • Forgetting the conditions (linearity, independence, equal spread, normal residuals) - the t-test is only valid when the regression assumptions hold.

Why This Formula Matters

A nonzero sample slope can appear from pure noise even when no real relationship exists, so describing the line isn't enough — you need a test that separates a genuine trend from random scatter. This is the step that lets you say 'there IS a linear relationship in the population,' which a single fitted line can never claim on its own. Recognizing it by "Am I testing whether the underlying population slope is nonzero (rather than just computing or describing the sample slope)?" — rather than by familiar numbers — is what lets a student tell it apart from lsrl and correlation test and two-sample t-test in a mixed problem set.

Frequently Asked Questions

What is the Inference for Regression formula?

Using hypothesis tests and confidence intervals to draw conclusions about the true population slope β1 of the linear relationship y=β0+β1x+ε, based on sample data.

How do you use the Inference for Regression formula?

You computed a sample regression line with slope b=2.3. But is the true population slope actually different from zero? Maybe there's really no linear relationship and you just got a slope by chance. The regression t-test asks: 'Is my sample slope far enough from zero that it's unlikely to have occurred by random variation alone?'

What do the symbols mean in the Inference for Regression formula?

b = sample slope, β1 = population slope, SEb = standard error of the slope, s = standard deviation of residuals, df=n−2.

Why is the Inference for Regression formula important in Math?

A nonzero sample slope can appear from pure noise even when no real relationship exists, so describing the line isn't enough — you need a test that separates a genuine trend from random scatter. This is the step that lets you say 'there IS a linear relationship in the population,' which a single fitted line can never claim on its own. Recognizing it by "Am I testing whether the underlying population slope is nonzero (rather than just computing or describing the sample slope)?" — rather than by familiar numbers — is what lets a student tell it apart from lsrl and correlation test and two-sample t-test in a mixed problem set.

What do students get wrong about Inference for Regression?

The procedure for inference for regression is the easy part; the trap is treating a nonzero sample slope as proof of a population relationship. Asking "Am I testing whether the underlying population slope is nonzero (rather than just computing or describing the sample slope)?" first is what keeps a correct-looking calculation from being attached to the wrong concept.

What should I learn before the Inference for Regression formula?

Before studying the Inference for Regression formula, you should understand: linear regression lsrl, residuals, r squared, hypothesis testing, confidence interval.