Bessel’s Correction

In the field of statistics, understanding the difference between population and sample statistics is crucial for accurate analysis. An important concept to understand is Besse’ls Correction. This note will explain the purpose and justification of this correction and it’s relationship to the concept of degrees of freedom.

What is Bessel’s Correction

Summary Statistics

Consider the sample and population statistics, summarised below:

PopulationSample
Variance
Mean

Comparing Population and Sample Statistics

The best estimate of the population mean from a sample is the mean of that sample, calculated in the exact same way, as a matter of fact, this fundamental assumption will later allow us to derive the gaussian distribution. The sample Variance, however, differs from the population variance in two important ways:

  1. The difference concerns and not
    • If we have a sample we likely do not know and can only use
      • If, however, we did know , then it would be more appropriate to use that and use the population formula to estimate
  2. The denominator is .

Bessel’s correction refers to the use of in the denominator instead of . This correction is applied to ensure that the sample variance is an unbiased estimator of the population variance. The sample variance will tend to underestimate the population variance When using the sample mean as an estimate of the population mean, Bessel’s Correction will adjust for this.

Bias

The sample mean is only an estimate of the population mean , this means that the sample variance will be a less accurate when calculated with the sample mean than it would be with the population mean.

Since the sample mean is calculated from the same data points used to estimate the variance, it will be closer to those data points. Hence, the calculated variance will be smaller than the population variance.

If the sample mean is equal to the population mean, the distance from each observation will, of course, be equal, however, if the sample mean moves away from the population mean so too must the observations and the observations will always be closer to the sample mean than the population mean, because the sample mean is the centre of those points. Because this distance between is less than sample variance will be less when using as an estimator of the variance and it must be adjusted for, this is what Bessel’s Correction does, the reason is used relates to degrees of freedom and this will be explained below.

Note that using the mean of the distances (albeit squared), is still the best estimate of the population variance, just like the mean value, an assumption that leads to the derivation of a Gaussian Distribution is that the mean value of a sample is the best estimate of the population mean.

In summary, Bessel’s Correction compensates for the fact that the sample mean is likely to be closer to the data points in the sample than the true population mean, causing the variance to be slightly underestimated when using n as the denominator.

A Note on Maximum Absolute Deviation

A common question that arises in statistics, tangential to this topic, is why the squared-distance is used rather than the absolute distance (i.e. ), this relates to the assumption mentioned before, i.e. that the sample mean is the best estimate of the population mean, this is discussed further in deriving the normal distribution.

Degrees of Freedom and Variance

Degrees of freedom (df) is a concept that is fundamental to statistics. It refers to the number of independent pieces of information or parameters that can be freely changed when estimating a statistical parameter. More specifically it reflects the number of values in the final calculation that can vary without violating any constraints.

When calculating the sample variance, the sample mean is used as an approximation of the population mean , this consumes one degree of freedom, to see this, consider:

if is already known, any one of the remaining observations can easily be calculated, e.g. to get :

Variance is concerned with the amount of variation between each point (observation), however, if one of the points is essentially fixed because it is a function of the others, then there is only points that are varying, and so the sample variance should only be calculated with respect to those n-1 points.

Conclusion

Bessel’s correction is necessary to accurately estimate the population variance when the population mean is not known. Understanding this correction and the connection to degrees of freedom lost by using sample mean is essential for accurate statistical analysis. By using Bessel’s correction, we can obtain reliable variance estimates and perform tests that rely on this value (e.g. t-test, ANOVA).

Further Proof

Variance in Terms of Expected Differences

Summing Variances

Multiplying Variance by a Constant}

Variance of a Sample Mean

Consider a sample of size , the population will have a variance and a mean , the sample will have a variance and a mean . Many different samples could be taken, so there is a distribution of values that could occur. The variance of sample means is given by , this is a part of the CLT and is shown here.

Bessel’s Correction

Suppose that the sample variance was given by the typical formula:

The expected value should be , in that sense it would be an un-biased estimator of the population mean.

The key here is to recognise that corresponds to a sample, by introducing we can solve for variance in terms of expected values and the sample mean. The sample introduces the bias so it is necessary to use the sample mean.

\subsection*{Expected Value Squared}

The expected value of is the mean value:

Expected Square Value

Recall the definition of Variance from earlier:

Applying this to the sample mean:

This step provides the key insight, if variance is taken on a population, there is no in the denominator

because the variance of a sample mean is different to the variance of a population statistic, the expected value of a squared sample mean is different to the expected value of a squared observation. The same from is the same one that leands to The expected value of a squared sample mean is less than the expected value of a squared observation, because the sample mean is going to be more central.

\section*{Solving The Expected Sample Variance}

This shows that the variance formula, applied to a sample is biased. In order to correct that bias define like so:

The expected value of this is: