IB Syllabus Requirements for Estimation & Confidence Intervals
4.12
Data collection, categorisation, reliability and validity
4.15
Sampling distributions and the central limit theorem
4.16
Confidence intervals for a population mean
4.12
DATA COLLECTION, CATEGORISATION, RELIABILITY AND VALIDITY
A survey collects information from a sample of individuals. The planned set of questions used to collect that information is a questionnaire. Before drafting the questions, identify the population and decide what information is actually needed. You should also know how each response will be analysed. Without that planning, it’s easy to collect a large amount of data that cannot be used.
Wording can affect validity:
One useful revision check is to ask whether two reasonable respondents could interpret the question differently. Also ask whether any answer has been made to sound more respectable. If the answer to either question is yes, change the wording. The contrast between a leading, vague question and a neutral, measurable version appears below.

A variable is a recorded characteristic that can take different values. When choosing from many possible variables, keep those that measure the research question, explain the response or help reveal plausible confounding effects. Don’t include a variable simply because it is available.
The data should match the intended population, place and time period. Check how sampling was carried out and how each variable was measured. Units and categories need to be consistent, while missing values or selection effects may distort the analysis. Social-media and marketing data show the risk clearly: even a very large dataset can exclude people who don’t use the platform. Recommendation algorithms may also influence the behaviour being measured.
Similar decisions arise in fieldwork across the sciences, geography, psychology, sport and health studies, as well as in business and design. Questionnaires work efficiently and can be standardised, though fixed responses may suppress nuance. Interviews allow clarification, but interviewer effects and inconsistent prompts can make comparisons less reliable. No method is universally strongest; it must fit the claim under investigation.
A chi-square table is a frequency table used to compare observed category counts with counts expected under a proposed model. Before numerical measurements can enter the table, they may need to be grouped into intervals. Choose category boundaries for a defensible reason, such as meaningful ranges in the context. They should not be selected afterwards to produce a preferred conclusion.
Each expected frequency should be greater than . When expected frequencies are too small, combine adjacent categories if the new grouping remains meaningful, then explain the revised categorisation. Some detail is lost when categories are combined, so a balance is needed between preserving information and meeting the conditions of the test.
In a chi-square goodness-of-fit test, parameters estimated from the same data reduce the degrees of freedom:
If parameters are supplied in the question rather than estimated from the sample, they do not reduce the degrees of freedom.
Reliability describes a measurement method that gives consistent results under comparable conditions. Validity describes a method that measures the intended characteristic well enough to support the proposed interpretation. Reliability concerns consistency. Validity asks whether the measurement is aimed at the right target. A method may be reliably wrong, while serious inconsistency can also limit validity.
| Method | What it checks | How it is used |
|---|---|---|
| Test-retest reliability | Stability over time | Give the same test to the same participants on separate occasions, then compare the results. |
| Parallel-forms reliability | Consistency across equivalent instruments | Give two versions designed to measure the same construct, then compare the results. |
| Content validity | Coverage of the intended domain | Check whether the items adequately represent all relevant parts of the characteristic being measured. |
| Criterion-related validity | Agreement with an appropriate external criterion | Compare the results with a trusted outcome or an established measure of the same or a closely related characteristic. |
High agreement in a reliability test does not establish validity on its own. A test may also contain relevant content yet still be weakened by biased sampling, unclear wording or an unsuitable external criterion.
4.15
SAMPLING DISTRIBUTIONS AND THE CENTRAL LIMIT THEOREM
A linear combination is a random variable made by multiplying random variables by constants and adding the results. Another constant may also be added:
Independent random variables are random variables where knowing the value of one does not change the probability distribution of the others. If independent normally distributed random variables are combined linearly, the result is also normally distributed. Their means and variances can differ. The required conditions are independence and normality.
The sample mean random variable is the arithmetic mean defined before a random sample has been observed:
For independent observations from the same normal population,
The sample mean is centred at the population mean and has standard deviation . This standard deviation is the standard error. It measures how much sample means vary between samples. A larger sample size reduces this variation, but it does not usually reduce the spread of the individual observations.
The central limit theorem states that the distribution of the sample mean approaches a normal distribution as the size of an independent sample increases, provided the population has a finite mean and variance. So, even if the population is not normal,
for sufficiently large . The population’s shape affects how large “sufficiently large” needs to be. A roughly symmetric population settles quickly; a strongly skewed population may need a larger sample. For this course, is treated as sufficient in examinations. When the population is already normal, the result for is exact for every sample size rather than an approximation.
A simulation shows the difference clearly. Repeatedly take samples from a non-normal population, calculate the mean of each sample and plot the results. As sample size increases, the distribution of these means becomes more bell-shaped and less spread out. It stays centred near .

This is why field studies can use results from many sampled individuals even when individual measurements are far from normal. The same result supports much of statistical estimation in the human sciences: averaging produces a distribution whose behaviour can be modelled predictably.
The theorem has a formal side and an empirical one. It can be established deductively from mathematical assumptions. Simulations and successful applications then provide empirical confirmation that those assumptions are useful. Applications cannot prove the theorem, though they can test whether its model describes a particular setting adequately.
For a finite population containing individuals, the number of distinct unordered samples of size without replacement is
4.16
CONFIDENCE INTERVALS FOR A POPULATION MEAN
A confidence interval is calculated from sample data using a method designed to capture an unknown population parameter in a stated proportion of repeated samples. In this case, the parameter is the mean of a normal population.
The confidence level gives the long-run proportion of intervals produced by the method that contain the true parameter. For example, a confidence level of describes how the method performs across repeated random samples. After a particular interval has been calculated, the population mean remains fixed, so we don’t say there is a probability that this fixed value lies within the interval. A suitable interpretation in context is, “We are confident that the population mean lies between the calculated endpoints.”
Intervals from many samples illustrate this repeated-sampling idea: most cross the fixed population mean, while a small proportion do not.

Use the normal distribution when the population standard deviation is known. The confidence interval takes the form
where is the observed sample mean (in the unit of the measured variable) and is the positive standard-normal critical value determined by the confidence level (dimensionless). The other symbols retain their earlier meanings.
The term after is the margin of error. Technology can provide the appropriate critical value and endpoints, but the interval must be stated and interpreted using the population and variable from the context. Since the population is normal, the method works for any sample size.
If is unknown, replace it with the sample standard deviation and use the -distribution:
Here, is the positive critical value for degrees of freedom (dimensionless), while is the sample standard deviation (in the unit of the measured variable). For this syllabus, use the -distribution whenever is unknown, regardless of sample size. The underlying population is assumed to be normal.
Because estimating adds uncertainty, the -distribution has heavier tails than the normal distribution. As the degrees of freedom increase, its critical values move closer to the corresponding normal critical values.
A wider interval gives a less precise estimate. Increasing the confidence level or observed variability makes the interval wider; increasing the sample size makes it narrower. The square-root effect matters here: a substantially narrower interval generally requires a much larger sample.
Confidence intervals put a numerical measure of uncertainty on forecasts and claims drawn from field-study data. Suppose an interval estimates mean waiting time. Both endpoints must then be reported in the time unit, and the interpretation must refer to the population mean—not the proportion of individual waiting times within the interval.
When two brands or treatments are compared, overlapping confidence intervals suggest that the claim “one is better on average” may not be strongly supported by the displayed estimates. Overlap alone, though, is not a formal test of equality. Its extent depends on the confidence level, sample sizes, variability and whether the samples are related. Confidence intervals can show the precision and practical importance of a claim, but they cannot make a contextual comparison certain.