Slide explaining the difference between residuals and errors in linear regression by comparing differences from the estimated line and true population line

What Is a Residual?

Take any individual $i$, and consider the observed values $y_i$ and $x_i.$ In our main example, these are their wages and experience.

Suppose now that we happen to lose or forget the value of their wage, $y_i.$ If we still have $x_i$ to hand, we can estimate $y_i$ from our “line of best fit”, or estimated model.

For an individual already within our sample, we call this a fitted value. Using our line of best fit, this is expressed as follows:

$$\hat y_i=\hat\beta_0+\hat\beta_1x_i.$$

The residual is precisely the difference between the actual value of their wage observed in the data and this fitted value:

$$\hat u_i=y_i-\hat y_i=y_i-(\hat\beta_0+\hat\beta_1x_i).$$

Geometrically, it appears as the vertical difference between the datapoint and our estimated line of best fit.

For worker 5 in the slide, the actual wage is $33{,}168.96$, while the fitted wage is $34{,}664.90.$ Hence,

$$\hat u_5=33{,}168.96-34{,}664.90=-1{,}495.94.$$

The negative residual tells us that this worker lies below our fitted line.

So far this may seem pointless, since we do know the true value of his wage – but in fact, residuals are central to linear regression, and very useful to investigate!


What Is an Error?

The error is quite similar, but it measures a different difference.

Recall that our original model is

$$Y=\beta_0+\beta_1X+U.$$

For individual $i$, this becomes:

$$y_i=\beta_0+\beta_1x_i+u_i.$$

The error $u_i$ is therefore the vertical difference by which this individual has been “pushed off” from the true population line – either up or down.

If we rearrange this equation, we can write their error term as follows:

$$u_i=y_i-(\beta_0+\beta_1x_i).$$


Comparing the Two

So, whilst the residual measures the vertical difference from the estimated line, $y=\hat\beta_0+\hat\beta_1x,$ the error measures the vertical difference from the true population line, $y=\beta_0+\beta_1x.$

Look again at the two formulae:

$$\hat u_i=y_i-(\hat\beta_0+\hat\beta_1x_i).$$

$$u_i=y_i-(\beta_0+\beta_1x_i).$$

In general, our estimates $\hat\beta_0$ and $\hat\beta_1$ will not be exactly the same as the true values $\beta_0$ and $\beta_1.$ So, the two comparison lines will generally not be the same either. Hence, the residual $\hat u_i$ and error $u_i$ will generally take different values.

Importantly, the true coefficients $\beta_0$ and $\beta_1$ are unknown. This means that the true error $u_i$ is not something we can actually work out the value of.

The residuals $\hat u_i$, however, depend only upon our estimates and the data, which we do know. This is their main “claim to fame”: they are observable, and we use them whenever we would like to learn about the true error terms, which sadly we can never observe directly.

We can think of them as knock-off versions of the error terms! In statistical language, we may sometimes think of $\hat u_i$ as an estimate of $u_i$ – hence adding the “hat” as usual.


Background

Found this useful?

Understanding Econometrics is completely free to use, and always will be.

If you found the site useful and would like to help me keep adding new material, please consider buying me a coffee!

Buy me a coffee ☕