Convergence in probability explained with intuitive definition and epsilon condition

What Is Convergence in Probability?

Consider a sequence of random variables

$$X_1,X_2,X_3,\dots$$

and another random variable $X$, which will act as our target.

The basic idea behind convergence in probability is that, for large $n$, the probability that $X_n$ and $X$ are “far apart” becomes very small.

To formalise this, we need to make clear what we mean by “far apart”.

Suppose, for now, that we choose to say that $X_n$ and $X$ are far apart whenever they differ by more than $0.1$:

$$|X_n-X|>0.1.$$

The probability of being far apart is then

$$\mathbb{P}(|X_n-X|>0.1).$$

If $X_n$ is converging towards $X$ in probability, we need this probability to become smaller and smaller as $n$ increases:

$$\lim_{n\to\infty}\mathbb{P}(|X_n-X|>0.1)=0.$$

But $0.1$ is just one possible choice of the distance that counts as “far apart”. Moreover, whether it is appropriate will depend on the situation.

If you’re designing some trainers and estimating average foot widths in centimetres, a difference of $0.1$ cm may already be small enough for your purposes. If you’re designing a precision component for a mission to Mars, being $0.1$ cm out might be far too much!

Our full definition, then, needs to work for any positive distance we choose.


The Full Definition

We use the Greek letter $\varepsilon$ to represent a chosen distance.

For any

$$\varepsilon>0,$$

the event

$$|X_n-X|>\varepsilon$$

means that $X_n$ and $X$ are more than $\varepsilon$ apart.

We say that $X_n$ converges in probability to $X$ if

$$\lim_{n\to\infty}\mathbb{P}(|X_n-X|>\varepsilon)=0$$

for every $\varepsilon>0$.

In words: whatever positive distance we choose, the probability that $X_n$ and $X$ are far apart according to this choice tends to zero as $n$ becomes large.

We use the following notation to express this:

$$X_n\overset{p}{\longrightarrow}X.$$


What Is Actually Converging?

The notation

$$X_n\overset{p}{\longrightarrow}X$$

looks rather like an ordinary limit, but there is an important difference.

Once we have fixed a value of $\varepsilon$, it is about the following sequence of probabilities:

$$\mathbb{P}(|X_1-X|>\varepsilon),\quad \mathbb{P}(|X_2-X|>\varepsilon),\quad \mathbb{P}(|X_3-X|>\varepsilon),\dots$$

Each term in this sequence is just a number between $0$ and $1$. Convergence in probability means precisely that this sequence of probabilities tends to zero.

In particular, the random variables $X_1, X_2, \dots $ themselves are not converging in the ordinary sense of a sequence of numbers – they are not numbers at all, but random variables!


How Does Convergence in Probability Compare with Other Notions?

The three main notions of convergence we are considering satisfy

$$\text{Almost Sure Convergence}\Rightarrow\text{Convergence in Probability}\Rightarrow\text{Convergence in Distribution}.$$

Almost sure convergence implies convergence in probability, and convergence in probability implies convergence in distribution.

The reverse implications do not hold in general.

There is one important special case. If the target random variable is simply a constant, say $X=c$, then convergence in distribution to $c$ is equivalent to convergence in probability to $c$.

In this case, if $X_n$ converges in distribution to $c$, it begins to “behave more like” the constant $c$ as $n$ increases. In particular, it becomes increasingly unlikely to be far away from $c$.


Example: The Weak Law of Large Numbers

An especially important example comes from sample averages.

Suppose

$$X_1,X_2,X_3,\dots$$

are IID random variables with

$$\mathbb{E}(X_i)=\mu$$

and finite variance

$$\operatorname{Var}(X_i)=\sigma^2.$$

Let

$$\bar{X}_n=\frac{X_1+X_2+\dots+X_n}{n}$$

be the average of the first $n$ observations.

The Weak Law of Large Numbers says that

$$\bar{X}_n\overset{p}{\longrightarrow}\mu.$$

In words, as the sample size becomes large, the probability that the sample mean and population mean are far apart becomes very small.

Using the full definition, for every $\varepsilon>0$,

$$\lim_{n\to\infty}\mathbb{P}(|\bar{X}_n-\mu|>\varepsilon)=0.$$


Background:

Found this useful?

Understanding Econometrics is completely free to use, and always will be.

If you found the site useful and would like to help me keep adding new material, please consider buying me a coffee!

Buy me a coffee ☕