Linear regression involves a lot of mathematics, but much of the subject revolves around one basic task: using data to estimate the unknown coefficients in a linear model.
We begin with a highly simplified example, designed to make absolutely clear what these coefficients mean, and how a line of best fit enables us to do so. From there, we introduce ordinary least squares (OLS) as the standard method for choosing this line.
The mathematics then becomes more involved. We investigate the assumptions under which OLS works well through the Gauss-Markov theorem, consider different functional forms for linear models, and ask how well our fitted models describe the data using the measure $R^2$.
Finally, we study the uncertainty surrounding our estimates of coefficients, and how to test claims about them. This leads to estimating variances and standard errors, and then to the $t$-test and $F$-test.
Although becoming quite involved, these topics are all developments of the same basic idea introduced at the beginning: using a sample of data and a line of best fit to learn about the unknown relationship between variables.
- Probability and Statistics Recap
-
Linear Regression Basics
The basic ideas behind linear regression, including linear models, lines of best fit, residuals, and ordinary least squares.
-
Gauss-Markov For Simple Linear Regression
The assumptions that make OLS work well in a simple linear regression, and precisely what they allow us to conclude about our estimators.
-
Functional Form
How logarithms and interaction terms change the interpretation of a linear regression model.
-
$R^2$ and Model Fit
How $R^2$ measures the proportion of variation in the dependent variable explained by a fitted regression model.
- Estimating Variance and Standard Errors [beta]
-
The $t$-test
How the regression $t$-test is used to test claims about individual coefficients in a linear regression model.
- The $F$-test [beta]