Chapter 13
STATISTICS
“Statistics may be rightly called the science of averages and their estimates.” – A.L. BOWLEY & A.L. BODDINGTON
13.1 Introduction
The discipline of statistics is fundamentally concerned with the systematic collection of data for predefined objectives. Through rigorous analysis and interpretation, informed decisions can be derived from such data. Prior coursework has introduced methodologies for visually depicting data, both graphically and through tabular formats. Such presentations serve to highlight intrinsic attributes or distinguishing characteristics of the dataset. Furthermore, we have explored techniques for identifying a singular value that epitomizes a given dataset, commonly referred to as a measure of central tendency. It is pertinent to recall that the mean (or arithmetic mean), median, and mode constitute the primary measures of central tendency. While a measure of central tendency offers an approximate indication of the data's focal point, a comprehensive understanding necessitates an assessment of the data's dispersion or its degree of concentration around this central value.
Karl Pearson (1857-1936)
Let us now examine the scores attained by two distinct batsmen over their preceding ten matches, presented as follows:
Batsman A : 30, 91, 0, 64, 42, 80, 30, 5, 117, 71
Batsman B : 53, 46, 48, 50, 53, 53, 58, 60, 57, 52
Evidently, the calculated mean and median for these datasets are as follows:
| Batsman A | Batsman B | |
|---|---|---|
| Mean | 53 | 53 |
| Median | 53 | 53 |
It is pertinent to recollect that the arithmetic mean of a dataset (symbolized by $\overline{x}$ ) is computed by partitioning the aggregate sum of all observations by the total count of those observations, specifically:
$
\overline{x} = \frac{1}{n} \sum_{i=1}^{n} x_i $
Concurrently, the median is ascertained by initially ordering the data points either in ascending or descending sequence, and subsequently applying the ensuing principle.
If the number of observations is odd, then the median is $\left(\frac{n + 1}{2}\right)^{\text{th}}$ observation.
If the number of observations is even, then median is the mean of $\left(\frac{n}{2}\right)^{\text{th}}$ and $\left(\frac{n}{2} + 1\right)^{\text{th}}$ observations.
It is observed that both Batsman A and Batsman B exhibit identical mean and median scores, specifically 53 runs. Does this equivalence imply comparable performance between the two athletes? Unequivocally not, given that Batsman A's scores demonstrate a considerable variability, spanning from a minimum of 0 to a maximum of 117. In contrast, Batsman B's scores display a much narrower range, extending only from 46 to 60.
To further illustrate this distinction, let us represent the aforementioned scores as discrete points on a numerical axis. The resulting visual representations are as follows:
For batsman A

For batsman B

Fig 13.1
Fig 13.2
It is discernible that the data points pertaining to Batsman B are tightly grouped, exhibiting a pronounced clustering effect around the measures of central tendency (i.e., the mean and median), whereas the data points for Batsman A are considerably more dispersed and widely distributed.
Consequently, relying solely on measures of central tendency does not furnish a comprehensive understanding of a given dataset. Variability constitutes an additional critical aspect necessitating statistical investigation. Analogous to central tendency metrics, the objective is to encapsulate variability within a singular numerical value. This singular value is termed a ‘measure of dispersion’. Within this Chapter, we will explore several significant measures of dispersion, alongside their computational methodologies for both ungrouped and grouped data.
13.2 Measures of Dispersion
The extent of dispersion or spread within a dataset is quantified by considering the individual observations and the specific type of central tendency measure employed. Key measures of dispersion include:
(i) Range, (ii) Quartile deviation, (iii) Mean deviation, (iv) Standard deviation.
This Chapter will delve into all these measures of dispersion, with the exception of the quartile deviation.
13.3 Range
As previously discussed, in the context of runs accumulated by two batsmen, A and B, an intuitive understanding of score variability was derived from the minimum and maximum runs observed in each sequence. To derive a singular numerical representation for this variability, one computes the disparity between the maximum and minimum values within each respective series. This calculated disparity is formally known as the 'Range' of the dataset.
For batsman A, the Range is calculated as 117 - 0 = 117, whereas for batsman B, it is 60 - 46 = 14. Evidently, the Range for A significantly exceeds that for B. Consequently, the scores achieved by batsman A exhibit greater scattering or dispersion, while those for batsman B are considerably more concentrated.
Hence, the Range of a series is defined as: Maximum value - Minimum value.
While the data range offers a preliminary indication of variability or spread, it does not convey information regarding the data's dispersion relative to a specific measure of central tendency. To address this limitation, alternative measures of variability are required. Evidently, any such measure must incorporate the difference (or deviation) of individual values from a chosen central tendency.
The principal measures of dispersion that rely on the deviations of observations from a central tendency include mean deviation and standard deviation. These will be examined in subsequent detail.
13.4 Mean Deviation
The divergence of an observation $x$ from a designated reference point $a$ is defined as the difference $x - a$. To quantify the spread of $x$ values around a central point $a$, we compute these deviations relative to $a$. While one might consider the mean of these deviations as a measure of absolute dispersion, a fundamental issue arises. Given that a measure of central tendency inherently falls within the range of the dataset's minimum and maximum values, some deviations will inevitably be negative while others are positive. Consequently, the aggregate sum of these deviations can be zero. Notably, the sum of deviations from the arithmetic mean $(\overline{x})$ is always precisely zero.
$ \text{Also} \quad \text{Mean of deviations} = \frac{\text{Sum of deviations}}{\text{Number of observations}} = \frac{0}{n} = 0 $
Therefore, calculating the mean of deviations around the arithmetic mean provides no meaningful insight for assessing dispersion.
To establish an appropriate metric for dispersion, it is essential to quantify the separation of each data point from a chosen central tendency or a specific value, denoted ‘$a$’. It is pertinent to recall that the absolute value of the difference between any two numbers geometrically represents their distance on a number line. Consequently, to derive a measure of dispersion relative to a fixed value ‘$a$’, one can compute the mean of the absolute magnitudes of the deviations from this central value. This particular average is termed the ‘mean deviation’. Therefore, the mean deviation around a central value ‘$a$’ is formally defined as the average of the absolute differences between the observations and ‘$a$’. The notation M.D. ($a$) designates the mean deviation from ‘$a$’. Hence,
$ \text{M.D.}(a) = \frac{\text{Sum of absolute values of deviations from } 'a'}{\text{Number of observations}}. $
Remark Mean deviation can be computed with respect to any measure of central tendency. Nevertheless, mean deviation calculated from the mean and from the median are the most frequently employed in statistical analyses.
We will now proceed to examine the methodology for calculating mean deviation with respect to the mean and the median for diverse data structures.
13.4.1 Mean deviation for ungrouped data
Given a set of $n$ observations as $x_{1}, x_{2}, x_{3}, \ldots, x_{n}$, the following procedural steps are executed to compute the mean deviation concerning either the mean or the median:
Step 1 Ascertain the specific measure of central tendency around which the mean deviation is to be calculated. Let this value be designated as ‘$a$’.
Step 2 Determine the individual deviation of each $x_{i}$ from $a$, expressed as $x_{1} - a, x_{2} - a, x_{3} - a, \ldots, x_{n} - a$.
Step 3 Obtain the absolute magnitudes of these deviations, which involves disregarding any negative sign, resulting in $|x_{1} - a|, |x_{2} - a|, |x_{3} - a|, \ldots, |x_{n} - a|$.
Step 4 Compute the arithmetic mean of these absolute deviations. This calculated mean represents the mean deviation about $a$, formally expressed as:
$ \text{M.D.}(a) = \frac{\sum_{i=1}^{n} |x_{i} - a|}{n} $
Consequently, M.D. $(\overline{x}) = \frac{1}{n} \sum_{i=1}^{n} |x_i - \overline{x}|$, where $\overline{x} = \text{Mean}$,
and M.D. $(\mathbf{M}) = \frac{1}{n} \sum_{i=1}^{n} |x_i - \mathbf{M}|$, where $\mathbf{M} = \text{Median}$.
Note: Within this chapter, the symbol M will consistently represent the median, unless explicitly specified otherwise. The subsequent examples will now demonstrate the application of the aforementioned method.
Example 1 Find the mean deviation about the mean for the following data:
6, 7, 10, 12, 13, 4, 8, 12
Solution: We will systematically follow a series of steps to arrive at the solution:
Step 1: The arithmetic mean of the provided data set is determined as:
$ \bar{x} = \frac{6 + 7 + 10 + 12 + 13 + 4 + 8 + 12}{8} = \frac{72}{8} = 9 $
Step 2: The individual deviations of each observation from the calculated mean, $\bar{x}$, specifically $x_i - \bar{x}$, are computed as:
6 - 9, 7 - 9, 10 - 9, 12 - 9, 13 - 9, 4 - 9, 8 - 9, 12 - 9,
or -3, -2, 1, 3, 4, -5, -1, 3
Step 3: The absolute magnitudes of these deviations, denoted as $\left|x_i - \bar{x}\right|$, are then:
3, 2, 1, 3, 4, 5, 1, 3
Step 4: Consequently, the mean deviation about the mean
is
calculated as: