Clastify logo
Clastify logo
Exam prep
Exemplars
Review
HOT

Statistics

Master IB Math AA Statistics with notes created by examiners and strictly aligned with the syllabus.

IB Syllabus Requirements for Statistics

4.1

Data, sampling and validity

4.2

Presentation of data

4.3

Measures of central tendency and dispersion

4.1

DATA, SAMPLING AND VALIDITY

What a data set is really telling you

Statistics is a branch of mathematics that collects, organises, analyses and interprets data so that we can make reasonable conclusions about a wider situation. Reasonable is the key word. Even a beautiful calculation gives poor evidence when the data itself is poor.

A population is a complete set of individuals or items about which information is wanted. A sample is a subset of a population chosen to provide data about that population. At SL, the data set provided is treated as the population for calculation purposes unless the question tells you otherwise. Real investigations usually rely on samples because measuring a whole population is often impractical.

A random sample is a sample selected by a process in which each member, or each sample of the required size, has a known and fair chance of being chosen. This protects against selecting only people or items that happen to be easy to reach, visible or convenient.

Types of data

Quantitative data is numerical data obtained by counting or measuring a quantity. There are two main forms. Discrete data is quantitative data that can take separated values, usually because it is counted. For example, the number of messages received in an hour is discrete. Continuous data is quantitative data that can take any value in an interval, usually because it is measured. In a mathematical model, time, length and mass are continuous, even though the measuring device will round them.

You should also recognise categorical data as data sorted into non-numerical groups. Although the syllabus focuses here on discrete and continuous data, identifying a categorical variable prevents you from calculating a mean that has no meaning.

Reliability, bias and missing values

A bias is a systematic tendency in the data collection process that makes some outcomes more likely to appear than they should. Dishonesty isn’t the only cause of bias. It may come from how a question is worded, the time of day when a survey takes place, who can access the internet or which cases are missing.

Before you trust a data set, work through a few questions. Who collected it, and how? Who was left out? Check for missing values and consider whether recording errors could have occurred. Missing data isn’t automatically useless, but the reason it is missing matters. Randomly missing entries may do little damage. By contrast, conclusions can be badly distorted if the missing cases are weaker responses, failed machines or absent participants.

Sampling techniques

A simple random sample is a sampling method in which every member of the population has an equal chance of selection. Statistically, this is often the cleanest method. The drawback is that it requires a complete list of the population.

A convenience sample is a sampling method that uses members of the population who are easiest to access. It’s quick and cheap, but usually weak. Asking whoever happens to be nearby may produce a sample that says more about your location than about the population.

A systematic sample is a sampling method that selects members at regular intervals from an ordered list after a starting point is chosen. The method is efficient. However, a hidden pattern in the list can create bias if it matches the sampling interval.

A quota sample is a sampling method that fills fixed category totals without random selection within those categories. The sample might match the population across visible categories, yet the non-random choices made within each category can still cause bias.

A stratified sample is a sampling method that divides the population into relevant subgroups and then randomly samples from each subgroup. These subgroups are called strata, which are categories formed by shared characteristics. Of these methods, stratified sampling is often the best choice when important groups within the population need to be represented.

Image

Sampling is where statistics meets ethics. With training, a misleading graph may be fairly easy to recognise. A non-representative sample is harder to detect because its numbers can look perfectly respectable. For that reason, major failures of prediction are often caused by poor sampling rather than faulty arithmetic. The same problem arises in biology, psychology, geography, economics, business management and many research methods: conclusions only travel as far as the sampling method allows.

Outliers: not automatically wrong

An outlier is a data value that lies more than 1.5×IQR1.5\times \operatorname{IQR} from the nearest quartile. Here

IQR=Q3Q1\operatorname{IQR}=Q_3-Q_1

Values below Q11.5IQRQ_1-1.5\operatorname{IQR} or above Q3+1.5IQRQ_3+1.5\operatorname{IQR} are therefore treated as outliers. In context, an outlier may be a genuine and important case. It could instead result from a measurement error, typing error or unusual sampling accident. Don’t remove it simply because it is inconvenient; first decide whether it is valid evidence.

This helps explain why statistics and mathematics can sometimes feel like different subjects. Mathematics may produce a perfectly correct calculation, while statistics asks whether that calculation honestly represents the world. Hidden sampling can make statistics very misleading, and deliberately using statistical methods to create a false impression is never intellectually honest.

4.2

PRESENTATION OF DATA

Frequency distributions

A frequency distribution is a table that records how often values or classes of values occur in a data set. With discrete data, each class may be an individual value. With continuous data, the classes are intervals.

In this syllabus, write class intervals as inequalities with no gaps. For example, a continuous class might be 10t<2010\le t<20, where tt is the measured variable in the units of the context. The left endpoint belongs to the interval, but the right endpoint doesn’t. As a result, every boundary value has exactly one place to go.

Grouping data makes a table clearer, though some detail is lost. You can no longer see the exact original values within each class. It’s a useful trade-off, not a free lunch.

Image

Histograms

A histogram is a graph that represents grouped numerical data using adjacent bars whose bases are class intervals. This course uses frequency histograms with equal class intervals, so each bar’s height gives the frequency. Frequency density histograms are not required here.

The touching bars show that the variable is numerical and ordered. Separate categories are better displayed in a bar chart with gaps. For classes such as 0t<50\le t<5, 5t<105\le t<10, 10t<1510\le t<15, and so on, the horizontal axis must show the intervals accurately, with no gaps.

You can also use a histogram to describe the distribution’s shape. A roughly symmetric distribution looks balanced. In a right-skewed distribution, the longer tail extends to the right; in a left-skewed distribution, it extends to the left. This skewness affects whether the mean gives a fair picture of the centre.

Cumulative frequency graphs

Cumulative frequency is a running total of frequencies up to a given value or class boundary. A cumulative frequency graph is a graph of cumulative frequency against the upper class boundary. This graph is particularly useful when estimating positions in an ordered data set.

Put cumulative frequency on the vertical axis and data values on the horizontal axis. For grouped continuous data, plot each total at the upper boundary of its class. Join the points using a smooth increasing curve or sensible line segments, depending on the expected style.

Image

To take a reading, move horizontally from the required cumulative frequency to the curve, then vertically down to the data axis. Read the median at N2\frac{N}{2}, where NN is the total number of data values, a unitless count. The lower quartile is at N4\frac{N}{4}, while the upper quartile is at 3N4\frac{3N}{4}.

A percentile is a value below which a given percentage of the data lies. Read the ppth percentile at p100N\frac{p}{100}N, where pp is the percentile number, a unitless percentage value, and NN is the total number of data values. Watch the wording in statements such as “60%60\% are taller than aa”. Here, 40%40\% are at or below aa, so the correct reading comes from 0.40N0.40N, not 0.60N0.60N.

The range is a measure of spread found by subtracting the smallest value from the largest value. The interquartile range is a measure of spread found by subtracting the lower quartile from the upper quartile. You can estimate both from a cumulative frequency graph. However, the interquartile range leaves out the most extreme quarters of the data, making it less affected by outliers.

Box and whisker diagrams

A box and whisker diagram is a one-dimensional statistical diagram showing the minimum, lower quartile, median, upper quartile and maximum of a distribution. Together, these values are often known as the five-number summary.

The box stretches from Q1Q_1 to Q3Q_3, and a line marks the median. Whiskers display the spread outside the box. Mark outliers with a cross instead of quietly extending a whisker to include them.

Box-plot summaries for two groups

StatisticAB
Minimum145
Q1Q_11810
Median2214
Q3Q_32626
Maximum3034
IQR816
Outliersnone×53, ×57\times 53,\ \times 57
Shaperoughly symmetricright-skewed

Box plots work well when comparing two distributions because centre and spread appear on the same scale. Use the medians to compare typical values and the interquartile ranges to compare middle spread. The ranges show total spread, while the symmetry of the boxes and whiskers indicates shape. When the box and whiskers are fairly balanced around the median, the data may plausibly be approximately normally distributed. If they are strongly unbalanced, be cautious.

Unlike a histogram, a box plot doesn’t show frequency directly. Each section represents about a quarter of the data, even when one section covers much more of the number line than another. That difference carries information; it isn’t a drawing error.

Different subjects use these representations because raw data by itself is not yet information. Tables and diagrams become useful when they support a valid comparison or decision, though they can hide detail as well. Always ask what has been summarised away.

4.3

MEASURES OF CENTRAL TENDENCY AND DISPERSION

Mean, median and mode

A measure of central tendency is a single value used to describe the centre or typical value of a data set. You need to know three of them: the mean, median and mode.

The mean is a measure of central tendency found by dividing the total of the data values by the number of data values. For ungrouped data, use

xˉ=xin\bar{x}=\frac{\sum x_i}{n}

The symbol \sum means “add all the terms of this type”. Technology can calculate the mean, but you should still understand what the number is doing.

The median is a measure of central tendency that gives the middle value when the data are arranged in order. With an odd number of values, take the central value. With an even number, find the mean of the two central values.

The mode is a measure of central tendency that gives the most frequently occurring value. There may be no mode, one mode, or more than one mode in a data set. For grouped data, the modal class is the class interval with the greatest frequency. Modal class questions in this course use equal class intervals only.

Estimating the mean from grouped data

Grouped data hide the exact values inside intervals, so you estimate the mean using mid-interval values. A mid-interval value is the value halfway between the lower and upper boundaries of a class interval.

For grouped data,

xˉfimifi\bar{x}\approx \frac{\sum f_i m_i}{\sum f_i}

Pay attention to the approximation sign. The calculation treats every value in a class as though it sits at the midpoint. That may be a sensible estimate, but it’s only exact if the data really balance that way.

Grouped frequency table for estimating a mean.

Class interval [marks]FrequencyMidpoint [marks]f×f\times midpoint [marks]
0–924.59.0
10–19514.572.5
20–29824.5196.0
30–39934.5310.5
40–49444.5178.0
Total28n/a766.0

Quartiles and discrete data

A quartile is a value that divides an ordered data set into quarters. The lower quartile is the 2525th percentile, the median is the 5050th percentile and the upper quartile is the 7575th percentile.

You can find quartiles for discrete data by hand or with technology, and several methods are accepted. As a result, two answers may differ slightly even when neither method is mathematically careless. In IB work, use technology when the question asks for it. Don’t panic if your calculator gives a quartile that differs from the hand method used in a textbook example.

Measures of dispersion

A measure of dispersion is a value that describes how spread out the data are. The range is quick and simple, though it only uses the two most extreme values. The interquartile range uses the middle half of the data, so it is usually more resistant to outliers.

The standard deviation is a measure of dispersion that describes the typical distance of data values from the mean. The variance is a measure of dispersion equal to the square of the standard deviation. If technology gives the standard deviation as ss, the variance is s2s^2. Here, ss is the standard deviation in the same units as the data, while s2s^2 is the variance in squared data units.

In this syllabus, use technology to calculate standard deviation and variance. Hand calculation may help with understanding, but it isn’t the efficient method in an exam. A calculator may show more than one value for standard deviation or variance. Follow the wording of the question and the course convention that a given SL data set is treated as the population unless the question says otherwise.

Effect of constant changes

Data transformations change centre and spread in predictable ways. Let cc be a constant in the same units as the data when it is added or subtracted. When it is used to multiply or divide, cc is a unitless scale factor.

Effect of constant changes on summary statistics

MeasureAdd ccSubtract ccMultiply by ccDivide by cc
Meanincreases by ccdecreases by ccmultiplied by ccdivided by cc
Medianincreases by ccdecreases by ccmultiplied by ccdivided by cc
Standard deviationunchangedunchangedmultiplied by c\|c\|divided by c\|c\|
Varianceunchangedunchangedmultiplied by c2c^2divided by c2c^2
Rangeunchangedunchangedmultiplied by c\|c\|divided by c\|c\|
Interquartile rangeunchangedunchangedmultiplied by c\|c\|divided by c\|c\|

Add cc to every data value and the mean and median both increase by cc. The standard deviation, variance, range and interquartile range stay unchanged. The data have shifted; they haven’t spread out.

Multiply every value by cc and the mean and median are also multiplied by cc. The standard deviation, range and interquartile range are multiplied by c|c|, whereas the variance is multiplied by c2c^2. Spread cannot be negative, so the absolute value is needed even when multiplication by a negative number reverses the order of the data.

Choosing and questioning summary measures

Outliers affect the mean, while the median and interquartile range are often more suitable for skewed data. Standard deviation fits naturally with the mean: both use every data value, and standard deviation measures spread around the mean.

People use these measures to compare variation in populations, crops, social indicators, reliability and maintenance, and economic indices. Sharing data internationally can be powerful, provided that definitions and measurement methods match. A “standard” statistic may still have more than one formula in use, so you need to know what your technology is calculating.

Statistics tends to privilege what can be measured, and that limitation is worth stating. A tidy mean or standard deviation can make a variable seem more objective than it really is. Use the tools, but keep checking whether the chosen numbers answer the actual question.

Were those notes helpful?

probability Probability