What Is Convergence in Distribution?
Suppose we have a sequence of random variables
$$X_1,X_2,X_3,\dots$$
and another random variable $Z$, which will act as the “limit” of the $X_n$ in the sense defined below.
Recall that the CDF of a random variable $X$ is
$$F_X(x)=\mathbb{P}(X\leq x).$$
Roughly, we say that $X_n$ converges in distribution to $Z$ if the CDFs of $X_n$ approach the CDF of $Z$ as $n$ becomes large.
To formalise this, we first fix a value of $x$. We then require that
$$\lim_{n\to\infty}F_{X_n}(x)=F_Z(x),$$
at every point $x$ where $F_Z$ is continuous.
If this holds for all such $x$, we write
$$X_n\overset{d}{\longrightarrow}Z.$$
So convergence in distribution means that, for large $n$, the distribution or overall probability behaviour of $X_n$ becomes increasingly similar to that of $Z$.
This is very similar to pointwise convergence of functions: we fix $x$ first, and then let $n$ increase. The only difference here is that we ignore points $x$ where $F_Z$ is not continuous. So, if $F_Z$ is continuous everywhere, convergence in distribution is precisely pointwise convergence of the CDFs $F_{X_n}$ to $F_Z$.
This idea contrasts with uniform convergence, where we would instead ask how “far apart” the whole functions $F_{X_n}$ and $F_Z$ are, focusing on the “worst case” across all values of $x$.
Example: t-Distributions Approaching the Normal Distribution
A useful example comes from the Student t-distribution.
Suppose, for $n\geq2$,
$$X_n\sim t_{n-1}.$$
As $n$ becomes large, the t-distribution becomes increasingly similar to the standard normal distribution.
If
$$Z\sim\mathcal{N}(0,1),$$
then
$$X_n\overset{d}{\longrightarrow}Z.$$
The CDF of $X_n$ approaches the standard normal CDF at every value of $x$.
There is no requirement here that the different random variables $X_1,X_2,\dots$ are independent. Convergence in distribution concerns their distributions, rather than how the random variables are related to one another.
The Central Limit Theorem
The Central Limit Theorem is one of the most important examples of convergence in distribution.
Suppose
$$X_1,X_2,X_3,\dots$$
are IID random variables with
$$\mathbb{E}(X_i)=\mu$$
and
$$\operatorname{Var}(X_i)=\sigma^2,$$
where $0<\sigma^2<\infty$.
Let
$$\bar{X}_n=\frac{X_1+\dots+X_n}{n}$$
be the mean of the first $n$ observations.
For large $n$, we can roughly think of the sample mean as having an approximately normal distribution:
$$\bar{X}_n\approx\mathcal{N}\left(\mu,\frac{\sigma^2}{n}\right).$$
This approximation applies very generally, even when the original $X_i$ are not normally distributed.
For a given sample size, the approximation will often be better when the original distribution is itself fairly close to normal. If the original distribution is very skewed or has heavy tails, we may need a much larger value of $n$ before the normal distribution gives a good approximation to the distribution of the sample mean. But if the original $X_i$ are themselves normally distributed, the sample mean is exactly normal for every $n$.
To state the Central Limit Theorem precisely, we use the concept of convergence in distribution.
We first standardise the sample mean by subtracting its mean and dividing by its standard deviation:
$$\left(\frac{\bar{X}_n-\mu}{\sigma/\sqrt{n}}\right).$$
This puts every term in the sequence on the same scale, allowing us to compare them with a fixed target distribution that does not depend on $n$.
We can now express the CLT rigorously as follows:
$$\left(\frac{\bar{X}_n-\mu}{\sigma/\sqrt{n}}\right)\overset{d}{\longrightarrow}Z,$$
where
$$Z\sim\mathcal{N}(0,1).$$
We sometimes express the same idea by writing
$$\left(\frac{\bar{X}_n-\mu}{\sigma/\sqrt{n}}\right)\overset{a}{\sim}\mathcal{N}(0,1),$$
to say that asymptotically, the sequence has a standard normal distribution.
Similar Distributions Do Not Mean Similar Values
Convergence in distribution concerns the distributions of the random variables. It does not say that the random variables themselves must take similar values.
For a simple example, roll two independent fair dice. Let $X$ be the result of the first die and $Y$ the result of the second.
Then $X$ and $Y$ have exactly the same distribution and exactly the same CDF. However, they are independent. Knowing that the first die shows $6$ does not make the second die especially likely to show a value close to $6$.
The same distinction matters for convergence in distribution. Saying that
$$X_n\overset{d}{\longrightarrow}Z$$
means that $X_n$ and $Z$ have increasingly similar distributions. It does not mean that $X_n$ itself becomes likely to be numerically close to $Z$ for large $n$.
How Does Convergence in Distribution Compare with Other Types?
There are three main notions of convergence that we will consider:
$$\text{Almost Sure Convergence}\Rightarrow\text{Convergence in Probability}\Rightarrow\text{Convergence in Distribution}.$$
Almost sure convergence is the strongest of these, while convergence in distribution is the weakest.
The reverse implications do not hold in general.
One exception is the following special case. If the limiting random variable is a constant, convergence in distribution and convergence in probability are equivalent.
Background:
See also:
Understanding Econometrics is completely free to use, and always will be.
If you found the site useful and would like to help me keep adding new material, please consider buying me a coffee! Buy me a coffee ☕