Clastify logo
Clastify logo
Exam prep
Exemplars
Review
HOT

Distributions

Master IB Math AI Distributions with notes created by examiners and strictly aligned with the syllabus.

IB Syllabus Requirements for Distributions

4.7

Discrete random variables and expected value

4.8

Binomial distribution

4.9

Normal distribution

4.14

Transformations, linear combinations and unbiased estimators

HL

4.7

DISCRETE RANDOM VARIABLES AND EXPECTED VALUE

Random variables and their distributions

A random variable is a function that gives a numerical value to each outcome of a random process. If its possible values form a finite or countable set, it is a discrete random variable. Typical examples include counts, scores and gains from a game.

A discrete probability distribution lists, or gives a rule for finding, the probability of every possible value of a discrete random variable. We write

P(X=x)P(X=x)

.

For the distribution to be valid:

  • every probability lies between 00 and 11;
  • the probabilities sum to 11.

The distribution may be given in a table or as a formula that applies to a stated set of values. If you’re given a formula, calculate each probability and check that the total is 11 before going any further. This quick check catches many transcription errors.

Discrete probability distribution with probabilities that sum to 1.

Possible value xxProbability P(X=x)P(X = x)
00.10
10.30
20.40
30.20
Total1.00

Expected value

The expected value is the probability-weighted long-run mean of a random variable. For a discrete distribution,

E(X)=i=1mxipiE(X)=\sum_{i=1}^{m}x_i p_i

The expected value doesn’t have to be a value that can actually occur. For example, an expected score of 2.42.4 is still meaningful even when only whole-number scores are possible.

In applications, make sure each outcome has the correct net value. In a paid game, gain means the winnings minus the amount paid—not simply the prize. A fair game is one in which the player’s expected net gain is zero. A positive expected gain favours the player, while a negative expected gain favours the organizer.

Games of chance make expectation easier to see, though they also show its limits. A casino can make a reliable long-run profit from a small advantage spread across many plays. An individual player’s short-run experience may be very different. So, describing a game as fair can involve more than calculation; access to information, unequal resources and the possibility of harm also matter. Probability may support gambling systems or economic models, but mathematics does not remove responsibility for how those models are used.

4.8

BINOMIAL DISTRIBUTION

Recognizing a binomial setting

A Bernoulli trial is a random trial with exactly two classified outcomes: success and failure. A binomial distribution is a discrete probability distribution used to model the number of successes in a fixed number of independent Bernoulli trials when the probability of success stays constant.

Write the model as

XB(n,p)X\sim B(n,p)

Before using this model, check four conditions: nn is fixed, there are two outcome categories, the trials are independent, and pp is the same for every trial.

Sampling with replacement often allows independence and a constant probability. Without replacement, the probability generally changes after each selection. A binomial model may therefore fail unless the population is sufficiently large compared with the sample.

Probabilities using technology

To find the probability of an exact number of successes, technology evaluates

P(X=r)P(X=r)

A binomial probability density command returns an exact probability; a cumulative command returns P(Xr)P(X\le r). Translate the wording first, then press the buttons:

  • “at most rr” means P(Xr)P(X\le r);
  • “fewer than rr” means P(Xr1)P(X\le r-1);
  • “more than rr” means 1P(Xr)1-P(X\le r);
  • “at least rr” means 1P(Xr1)1-P(X\le r-1).

Mean and variance

For XB(n,p)X\sim B(n,p),

E(X)=npE(X)=np

and

Var(X)=np(1p)\operatorname{Var}(X)=np(1-p)

A success count is dimensionless, so its numerical variance is dimensionless too. The expected number of successes is the number of trials multiplied by the probability of success. You don’t need formal proofs of these results; focus on deciding when the model applies and using the results correctly.

Image

Binomial coefficients form the triangular pattern commonly linked with Pascal. However, Yang Hui recorded the pattern in China centuries earlier. The name attached to a result can reflect how knowledge travelled rather than who discovered it first.

Choose a model from evidence about its assumptions, not simply because a graph looks plausible. Binomial hypothesis testing is a natural later application, but here it is enrichment rather than part of the required calculations in this section.

4.9

NORMAL DISTRIBUTION

The normal model

A normal distribution is a continuous probability distribution with a symmetric, bell-shaped density curve. Its mean and variance determine the curve completely. We write

XN(μ,σ2)X\sim N(\mu,\sigma^2)

Because the curve is symmetric about μ\mu, the mean, median and mode coincide. Probability is shown by area under the curve, and the total area is 11. In both directions, the curve approaches the horizontal axis but never meets it. Changing μ\mu moves the curve, while changing σ\sigma alters its spread.

Image

For a continuous variable, the area at any single point is zero. Strict and inclusive inequalities therefore give the same normal probability. For example, P(X<x)=P(Xx)P(X< x)=P(X\le x).

The empirical rule

Standard deviation provides a quick picture of how values concentrate around the mean:

  • approximately 68%68\% of values lie from μσ\mu-\sigma to μ+σ\mu+\sigma;
  • approximately 95%95\% lie from μ2σ\mu-2\sigma to μ+2σ\mu+2\sigma;
  • approximately 99.7%99.7\% lie from μ3σ\mu-3\sigma to μ+3σ\mu+3\sigma.

These percentages are approximate features of the model. They do not prove that observed data are normal.

Empirical rule bands for a normal distribution.

BandInterval for XXArea in bandCumulative area
Central bandμσ\mu-\sigma to μ+σ\mu+\sigma68%68%
Next bandμ2σ\mu-2\sigma to μσ\mu-\sigma and μ+σ\mu+\sigma to μ+2σ\mu+2\sigma27% total95%
Outer bandμ3σ\mu-3\sigma to μ2σ\mu-2\sigma and μ+2σ\mu+2\sigma to μ+3σ\mu+3\sigma4.7% total99.7%
TailsX<μ3σX<\mu-3\sigma or X>μ+3σX>\mu+3\sigma0.3% totalabout 100%

Normal patterns often arise when many small, roughly independent effects influence a measurement. Examples appear in experimental science and psychology, as well as environmental measurement. Symmetry still needs to be checked. One normal curve may not represent bounded, strongly skewed or multimodal data well.

Normal and inverse normal calculations

A normal cumulative-distribution command calculates an area when the boundaries, μ\mu and σ\sigma are known. First sketch the curve, mark the boundary and shade the required region. This makes the calculator output easier to interpret. Enter both bounds for an interval. For a tail, either use an appropriate sufficiently extreme calculator bound or subtract a cumulative probability from 11.

An inverse normal calculation finds a boundary value from a given cumulative area. Technology needs the cumulative area to the left, along with the supplied mean and standard deviation. If a right-tail area is given, convert it to a left-tail area first. Use technology directly here; there is no need to transform to a standardized normal variable.

The normal curve developed through work that included de Moivre’s. Quetelet later used averages to form the idea of a representative person. This history carries a warning: treating variation around an average as deficiency can produce unsafe scientific or social conclusions. A model is trustworthy only when its assumptions and chosen variables suit its purpose. Choosing which observations to include or exclude is therefore a mathematical and ethical decision, not neutral housekeeping.

4.14

TRANSFORMATIONS, LINEAR COMBINATIONS AND UNBIASED ESTIMATORS

HL

Linear transformations

A linear transformation of a random variable creates a new random variable by multiplying the original variable by a constant, then adding another constant. Write

Y=aX+bY=aX+b

The expected value and variance are

E(aX+b)=aE(X)+bE(aX+b)=aE(X)+b

and

Var(aX+b)=a2Var(X)\operatorname{Var}(aX+b)=a^2\operatorname{Var}(X)

Adding bb shifts every value by the same amount. It changes the mean, but not the spread. Multiplying by aa scales the standard deviation by a|a| and the variance by a2a^2. A negative scale reverses the order of the values, although the variance remains non-negative. The general computational formula for variance is not required here.

Image

Linear combinations

A linear combination of random variables forms a random variable by adding constant multiples of other random variables. Let

L=c+j=1najXjL=c+\sum_{j=1}^{n}a_jX_j

Expectation is always linear:

E(L)=c+j=1najE(Xj)E(L)=c+\sum_{j=1}^{n}a_jE(X_j)

When the random variables are independent, their variances combine as

Var(L)=j=1naj2Var(Xj)\operatorname{Var}(L)=\sum_{j=1}^{n}a_j^2\operatorname{Var}(X_j)

This variance result depends on independence. Notice, too, that a negative sign becomes positive when its coefficient is squared. Independent variation does not cancel simply because one quantity is subtracted.

Estimating a population mean

An estimator is a statistic calculated from sample data to estimate an unknown population parameter. An unbiased estimator has an expected value equal to the parameter it estimates over repeated random samples.

The sample mean is

xˉ=1ni=1nxi\bar{x}=\frac{1}{n}\sum_{i=1}^{n}x_i

It is unbiased because

E(Xˉ)=μE(\bar{X})=\mu

A formal demonstration of this equality can help explain the result, but it is not required.

Estimating a population variance

Dividing the sum of squared deviations by nn tends to underestimate the population variance. The corrected sample variance is

sn12=nn1sn2=i=1kfi(xixˉ)2n1,n=i=1kfis_{n-1}^2=\frac{n}{n-1}s_n^2 =\sum_{i=1}^{k}\frac{f_i(x_i-\bar{x})^2}{n-1}, \qquad n=\sum_{i=1}^{k}f_i

Therefore,

E(sn12)=σ2E(s_{n-1}^2)=\sigma^2

The demonstrations of the two unbiasedness results are not examined. You need to understand why the divisor n1n-1 appears and be able to distinguish the sample statistic from the unknown population parameter.

Unbiased doesn’t automatically mean best. One unbiased estimator may vary greatly from sample to sample, whereas a slightly biased estimator may be more stable. When judging an estimator, consider its systematic error and variability, along with the purpose of the estimate.

4.17

POISSON DISTRIBUTION

HL

The Poisson model

A Poisson distribution is a discrete probability distribution used to model how many independent events occur in a fixed interval when the average rate is uniform. It is written

XPo(λ)X\sim \operatorname{Po}(\lambda)

An interval might represent time, distance, area or another continuous exposure measure.

The model is appropriate when:

  • events occur independently;
  • the average event rate remains uniform over the interval of interest.

For exactly rr events, the probability is

P(X=r)=eλλrr!P(X=r)=\frac{e^{-\lambda}\lambda^r}{r!}

In practice, technology is used to find Poisson probabilities.

Both the mean and the variance are numerically equal to the parameter:

E(X)=λ,Var(X)=λE(X)=\lambda, \qquad \operatorname{Var}(X)=\lambda

Counts are dimensionless in SI, so there is no unit conflict in this numerical equality. Formal proofs of these results are not required.

Image

When a rate is given for one interval but the question uses another, scale λ\lambda in direct proportion to the interval length. That calculation assumes the average rate is genuinely uniform. A clear rush-hour pattern or seasonal cycle would weaken the assumption.

Adding independent Poisson variables

Suppose X1Po(λ1)X_1\sim\operatorname{Po}(\lambda_1) and X2Po(λ2)X_2\sim\operatorname{Po}(\lambda_2) are independent. The variables X1X_1 and X2X_2 are event counts for the two sources and are dimensionless, while λ1\lambda_1 and λ2\lambda_2 are their respective expected counts and are also dimensionless. Then

X1+X2Po(λ1+λ2)X_1+X_2\sim\operatorname{Po}(\lambda_1+\lambda_2)

So independent streams of events can be combined. If they aren’t independent, this result cannot simply be assumed.

Selecting a distribution

Let the context determine the distribution rather than choosing whichever calculator menu feels most familiar.

Comparison of binomial, Poisson and normal models

ModelVariable typeDefining conditionsParametersTypical contexts
BinomialDiscrete count of successesFixed number of independent trials with constant success probability ppnn, ppSuccesses in repeated trials
PoissonDiscrete count of eventsEvents independent; average rate uniform over a fixed intervalλ\lambdaCounts over time, distance or area
NormalContinuous measurementRoughly symmetric data described by a mean and standard deviationμ\mu, σ\sigmaContinuous measurements with symmetric spread

Choose a binomial model for successes in a fixed number of independent trials when the success probability is constant. A Poisson model suits an event count over a continuous interval, provided the events are independent and the average rate is uniform. For a continuous, roughly symmetric measurement determined by a mean and standard deviation, use a normal model.

Poisson models may describe incoming requests, vehicles passing a sensor, rare changes in genetic material, unscheduled arrivals at a service point or printing defects, as long as the assumptions are credible. That usefulness doesn’t make the models literally true. They can support planning in science, public services or communication systems, but confidence in their conclusions should decrease when independence or rate stability is doubtful. Mathematical models organize evidence across many areas of knowledge; contextual judgement is still needed to decide which simplifications are acceptable.

Were those notes helpful?

descriptive-statistics Descriptive Statistics

estimation-and-confidence-intervals Estimation & Confidence Intervals