IB Syllabus Requirements for Distributions
4.7
Discrete random variables and expected value
4.8
Binomial distribution
4.9
Normal distribution
4.12
Standardization of normal variables
4.7
DISCRETE RANDOM VARIABLES AND EXPECTED VALUE
A random variable gets its value from the outcome of a random process. A discrete random variable can take only separated, countable values, such as , rather than every value in an interval.
We normally write for the random variable and for a particular value it can take. So is the probability that takes the value .
A probability distribution is a rule, table or formula that pairs each possible value of a random variable with its probability. In a discrete distribution, each probability must lie between and . All the probabilities must add to :
A probability distribution may appear as a table or as a formula, for example for . Either way, start by identifying the allowed values of . Calculate each probability if necessary, then check that the total probability is .
Discrete distribution and expected value calculation.
| 0 | 0.10 | 0.00 |
| 1 | 0.30 | 0.30 |
| 2 | 0.40 | 0.80 |
| 3 | 0.20 | 0.60 |
| Total | 1.00 | 1.70 |
The expected value is the long-run mean value of a random variable when the same random process is repeated many times. It doesn’t have to be a value that can actually occur. For instance, the expected score on one fair die is , even though no face marked exists.
For a discrete random variable,
Read the calculation as value times probability, then add the results. Don’t divide by the number of entries afterwards. The probabilities already perform the averaging, and they total .
Before calculating in an application, define clearly. If represents a gain, positive values usually show that the player wins money, while negative values show a loss. A fair game has an expected gain of zero, so neither side gains a long-run advantage:
Be a little sceptical here. Clear rules can make a game feel fair, but mathematical fairness requires zero expected gain. Casinos and lotteries usually use games in which the player has negative expected value. The rules may be transparent, but the long-run balance isn’t neutral.
4.8
BINOMIAL DISTRIBUTION
A binomial distribution is a discrete probability model that counts successes across a fixed number of independent trials. Each trial must have exactly two outcomes, and the probability of success must stay constant.
Use a binomial model only if all four conditions hold:
If counts the successes, write
The probability of failure is often denoted by , where
The phrase number of is often a clue, but it isn’t enough on its own. If the probability changes from one trial to the next, the model isn’t binomial unless the question tells you to assume independence and a constant probability.

For exactly successes, the probability is
In examinations, use technology to find binomial probabilities. Choose the exact-value function for a single value such as and the cumulative function for a range such as . For wording such as at least one, using the complement is often cleaner:
For ,
Its variance is
A formal proof of these results isn’t required here.
The standard deviation is therefore
The binomial coefficients also appear in the triangular pattern often linked with Pascal, although other mathematical traditions knew the pattern earlier. This reminds us that a mathematical discovery can have a wider history than the name attached to it.
A TOK question sits behind binomial work: what makes a model acceptable? A calculator’s ability to produce an answer doesn’t make the binomial model true. The model is appropriate when its assumptions reasonably approximate the context. If the trials affect one another or the probability of success changes, the calculation may be precise while the model itself is invalid.
4.9
NORMAL DISTRIBUTION
A normal distribution is a continuous probability model. Its graph is a symmetric, bell-shaped curve centred at the mean. When the context supports the assumption, it often models naturally varying measurements well—for example, biological measurements, manufactured dimensions or test scores.
We write
For a continuous random variable, probability comes from the area under the curve rather than the curve’s height at a point. Thus, is the area between two vertical lines, whereas is for one exact value.

The normal curve is symmetric about . The mean, median and mode therefore sit together at the centre, where the curve reaches its highest point. Its points of inflection are at and .
You should know these empirical percentages:
Use these percentages to interpret scale and catch unreasonable answers. A calculated probability that makes a value within one standard deviation seem extremely rare signals that something has gone wrong.
Use technology for normal probability and inverse normal calculations. In a probability calculation, you start with values of the variable and find an area. An inverse normal calculation works the other way round: start with an area and find a value of the variable.
At this stage, inverse normal questions give you the mean and standard deviation. You don’t need to transform to a standardized normal variable here. Draw or picture the curve, mark the mean and shade the required area. Then let the technology complete the calculation.
The normal distribution appears naturally in many settings, particularly when many small independent sources of variation combine. It can still be misused. Not every bell-ish histogram is normal, and distance from an average shouldn’t be used to judge every human measurement. The model can help us make predictions, but those predictions are only valid when the assumptions behind the model and the supporting data are valid.
4.12
STANDARDIZATION OF NORMAL VARIABLES
A standardized value shows how many standard deviations a value lies from the mean. For a normally distributed variable, this is called a z-value.
Use the standardization formula
If is positive, the value lies above the mean; if is negative, it lies below the mean. For example, means two standard deviations above the mean. It does not mean that the original value is .

If , standardizing it gives a variable that follows the standard normal distribution:
The standardization formula can also be rearranged:
I use this form most often when an inverse normal calculation gives me a z-value and I need to convert it back into the original units of the problem.
If the mean or standard deviation is unknown, use technology to find the appropriate z-value from the probability. Then form an equation using
When is unknown, rearrange to
When is unknown and , rearrange to
Use technology to find probabilities and variable values. The z-value then shows how the unknown parameter fits the probability statement. This topic links descriptive statistics with probability modelling: standardizing removes the original units, so different normal variables can be compared on the same scale.
4.14
FURTHER PROPERTIES OF RANDOM VARIABLES
The variance of a random variable measures spread by finding the expected squared distance from the mean. For a discrete random variable with mean ,
Here, is the variance of , is a possible value of , is the mean of , and is the probability of that value.
An equivalent form, which is often more convenient, is
where is the expected value of the square of .
To calculate , square each value before multiplying by its probability:
The standard deviation measures spread and is equal to the square root of the variance:
Variance can’t be negative. If your answer is negative, don’t take its square root and continue. Check the arithmetic, paying particular attention to the squaring in and .
A continuous random variable can take any value in an interval. A probability density function is a non-negative function with total area over the real line; the area over an interval gives its probability.
For a density function ,
where is the probability density function and is the continuous variable.
For an interval,
Here, and are the interval endpoints. With continuous variables, using or at the endpoints makes no difference because the probability of a single exact value is .
Piecewise functions appear often. Work with each piece on its own interval, splitting the integral whenever the formula for changes.

For a continuous random variable, the mean is
where is the mean of .
Its second moment is
The variance is still
As before, the standard deviation is the square root of the variance.
A mode of a continuous random variable is a value of where the probability density function reaches a maximum. If the density has several equally high peaks, there may be more than one mode.
A median of a continuous random variable is a value that divides the total area under the density curve into two equal halves:
where is the median.

Expected value often helps with decision-making. If represents gain, then is favourable in the long run and is unfavourable. When , the situation is fair in the strict mathematical sense.
This interpretation can be used in games, insurance and business decisions, though it doesn’t tell the whole story. A risky option may have a positive expected value but still be unacceptable to someone who can’t afford the possible loss. Mathematics clarifies the trade-off, but judgement is still needed.
A linear transformation of a random variable creates a new random variable by multiplying by a constant and then adding a constant. If
where is the transformed random variable, is the multiplier, is the original random variable, and is the added constant, then
and
Adding shifts every value. It changes the mean, but not the spread. Multiplying by stretches or reflects the distribution, so the variance is multiplied by . Don’t miss the square: multiplying by changes the variance by a factor of , not by .
In some investigations, another discrete distribution may provide a better model. The relationship between spread measures such as interquartile range and standard deviation also depends on the distribution used. A calculation is only as trustworthy as the model choice behind it.