Clastify logo
Clastify logo
Exam prep
Exemplars
Review
HOT

Probability

Master IB Math AA Probability with notes created by examiners and strictly aligned with the syllabus.

IB Syllabus Requirements for Probability

4.5

Probability basics, sample spaces and complementary events

4.6

Diagrams, combined events, conditional probability and independence

4.11

Formal conditional probability and testing for independence

4.13

Bayes' theorem for up to three events

HL

4.5

PROBABILITY BASICS, SAMPLE SPACES AND COMPLEMENTARY EVENTS

The language of probability

A trial is a procedure that can be repeated and gives one result. Examples include rolling a die once, selecting one card, or asking one student whether they travelled by bus. An outcome is one possible result of that trial. A sample space is the set of every possible outcome. In this syllabus, it is written as UU, where UU is the sample space being discussed and has no unit.

An event is a subset of the sample space containing the outcomes of interest. If AA is an event, n(A)n(A) gives the number of outcomes in event AA. It is a count, so it has no unit. In the same way, n(U)n(U) gives the number of outcomes in the sample space UU and is also a count with no unit.

When outcomes are equally likely, they all have the same chance of occurring. In this common special case,

P(A)=n(A)n(U)P(A)=\frac{n(A)}{n(U)}

The words equally likely matter. Counting sectors won’t give the probability if a spinner has unequal sectors. For a fair die, however, counting faces works. Probability lies between 00 and 11: 00 represents an impossible event, 11 represents a certain event, and values between them measure risk or likelihood.

Representing sample spaces

You can show a sample space as a list, a set, a table, or a grid. A list is enough for one die: U={1,2,3,4,5,6}U=\{1,2,3,4,5,6\}. With two trials, a table often helps prevent missed cases by making you account for every pair of outcomes.

Two fair dice sample space with at least one 6 highlighted.

T1T_1T2=1T_2=1T2=2T_2=2T2=3T_2=3T2=4T_2=4T2=5T_2=5T2=6T_2=6
1(1,1)(1,1)(1,2)(1,2)(1,3)(1,3)(1,4)(1,4)(1,5)(1,5)(1,6)\mathbf{(1,6)}
2(2,1)(2,1)(2,2)(2,2)(2,3)(2,3)(2,4)(2,4)(2,5)(2,5)(2,6)\mathbf{(2,6)}
3(3,1)(3,1)(3,2)(3,2)(3,3)(3,3)(3,4)(3,4)(3,5)(3,5)(3,6)\mathbf{(3,6)}
4(4,1)(4,1)(4,2)(4,2)(4,3)(4,3)(4,4)(4,4)(4,5)(4,5)(4,6)\mathbf{(4,6)}
5(5,1)(5,1)(5,2)(5,2)(5,3)(5,3)(5,4)(5,4)(5,5)(5,5)(5,6)\mathbf{(5,6)}
6(6,1)\mathbf{(6,1)}(6,2)\mathbf{(6,2)}(6,3)\mathbf{(6,3)}(6,4)\mathbf{(6,4)}(6,5)\mathbf{(6,5)}(6,6)\mathbf{(6,6)}

A relative frequency is the proportion observed after repeated trials. If an event occurs hh times in NN trials, its relative frequency is

hN\frac{h}{N}

.

This is experimental probability because it comes from data. Theoretical probability comes from a model, such as the assumption that a coin is fair. After many repetitions, the relative frequency may settle near the theoretical probability, but it remains an approximation rather than a guarantee.

That makes simulations useful. A calculator or spreadsheet can simulate a random process thousands of times very quickly, so the connection between theory and observation becomes visible. Simulations also show that probability isn’t certainty: a 0.10.1 chance does not mean exactly one occurrence in every ten trials.

Complementary events

The complement of an event contains every outcome in the sample space that is not in the original event. The complement of AA is written AA', where AA' is the event that AA does not occur.

Either AA occurs or AA' occurs, but they cannot both occur. Therefore,

P(A)+P(A)=1P(A)+P(A')=1

This is often the quickest method for phrases such as "at least one". Rather than listing every possible way to get at least one success, calculate 11- the probability of getting none.

Image

Expected number of occurrences

The expected number of occurrences is the long-run average count that a probability model predicts over a fixed number of trials. Suppose an event has probability pp, where pp is the probability of the event occurring in one trial and is dimensionless. If there are NN trials, where NN is the number of trials and is a count with no unit, then

Nexpected=NpN_{\text{expected}}=Np

For example, suppose a school expects a 0.080.08 probability that a student will be late on a wet morning. Among 250250 students, the expected number who are late is 250×0.08=20250\times 0.08=20. This does not require exactly 2020 students to be late; an expected value is an average produced by the model. That distinction matters when probability is used in insurance, planning, genetics, physics, or risk decisions, where people may perceive danger differently from how the numbers describe it.

4.6

DIAGRAMS, COMBINED EVENTS, CONDITIONAL PROBABILITY AND INDEPENDENCE

Choosing the right representation

Probability questions are easier once you can see their structure. A Venn diagram is a set diagram showing events as regions within a sample space. Use one for overlaps such as AA and BB, or for wording such as "or", "and", and "not".

A tree diagram is a branching diagram used to show a sequence of trials. It is particularly useful when probabilities change after the first outcome, especially in selection without replacement. A sample space diagram is a grid or table listing paired outcomes. It works well for two dice or two spinners, and for two independent selections where each pair can be counted.

When to use each probability diagram or table

RepresentationBest forTypical wording / structure
Venn diagramOverlaps of events"or", "and", "not"; sets inside a sample space
Tree diagramSequences of trialsSteps that change after the first outcome; with or without replacement
Sample space diagramPaired outcomesTwo dice, two spinners, or two selections where each pair can be counted
Outcomes tableListing all outcomesA grid of results to compare paired outcomes directly

Combined events and the meaning of "or"

Let BB be a second event in the same sample space. It is a subset of UU and has no unit. The event ABA\cup B is the union of events AA and BB, meaning AA or BB or both. In everyday English, "or" is often understood as exclusive. In probability, it is normally inclusive unless the question states otherwise.

The general addition rule is

P(AB)=P(A)+P(B)P(AB)P(A\cup B)=P(A)+P(B)-P(A\cap B)

Here, P(B)P(B) is the probability that event BB occurs. The event ABA\cap B means that both AA and BB occur, while P(AB)P(A\cap B) is the probability of that intersection. All these probabilities are dimensionless.

The overlap is already counted once in P(A)P(A) and again in P(B)P(B), so P(AB)P(A\cap B) must be subtracted. This is the safe formula to use unless the question gives a reason to simplify it.

Image

Mutually exclusive events

Two events are mutually exclusive events when they cannot occur in the same trial. In symbols,

P(AB)=0P(A\cap B)=0

For mutually exclusive events, the general addition rule becomes

P(AB)=P(A)+P(B)P(A\cup B)=P(A)+P(B)

Don’t confuse mutually exclusive events with independent ones. If two non-impossible events are mutually exclusive, the occurrence of one tells you that the other did not occur. They are therefore not independent.

Conditional probability

A conditional probability is a probability calculated after some information is known. The notation P(AB)P(A\mid B) means the probability that AA occurs given that BB has occurred, and P(AB)P(A\mid B) is dimensionless. Its formal definition is

P(AB)=P(AB)P(B)P(A\mid B)=\frac{P(A\cap B)}{P(B)}

provided P(B)0P(B)\ne 0. Rearranging gives

P(AB)=P(B)P(AB)P(A\cap B)=P(B)P(A\mid B)

The word "given" is a useful warning light. Once you know that BB has happened, the sample space is no longer all of UU; it is now the part inside BB. On a Venn diagram, compare the overlap ABA\cap B with the whole of BB, not with the entire rectangle.

Image

With and without replacement

In selection problems, with replacement means that an item is returned before the next selection, restoring the composition of the collection. Without replacement means that the item is not returned, so later probabilities usually change.

A tree diagram shows this difference clearly. Multiply branch probabilities along a path. When different paths satisfy the event, add their probabilities. If the tree includes every possible outcome, its terminal probabilities should add to 11. This provides a useful built-in check.

Image

Independent events

Two events are independent events when knowing that one has occurred does not change the probability of the other. For independent events,

P(AB)=P(A)P(B)P(A\cap B)=P(A)P(B)

Don’t assume this formula applies simply because two events appear unrelated. Independence is a condition to check, or information to use when the question states that the events are independent. The distinction matters in gambling contexts: probability can calculate risk clearly, but using it to exploit players raises an ethical question rather than a purely mathematical one.

4.11

FORMAL CONDITIONAL PROBABILITY AND TESTING FOR INDEPENDENCE

The formal definition revisited

The formal definition of conditional probability isn’t new, but this syllabus point requires you to use it deliberately:

P(AB)=P(AB)P(B)P(A\mid B)=\frac{P(A\cap B)}{P(B)}

with P(B)0P(B)\ne 0.

For tree diagrams, the rearranged form is often more useful:

P(AB)=P(B)P(AB)P(A\cap B)=P(B)P(A\mid B)

A complete path probability comes from multiplying the first-branch probability by the conditional probability on the second branch.

If the given event is BB', use the companion version:

P(AB)=P(B)P(AB)P(A\cap B')=P(B')P(A\mid B')

Image

Independence using conditional probability

There are two equivalent ways to test for independence. The multiplication test is

P(AB)=P(A)P(B)P(A\cap B)=P(A)P(B)

The conditional-probability test is

P(AB)=P(A)=P(AB)P(A\mid B)=P(A)=P(A\mid B')

If AA is independent of whether BB occurs, its probability stays the same inside BB, outside BB, and across the whole sample space. In applications such as medical studies or risk-factor analysis, this tests mathematically whether one classification appears to affect another.

Testing for independence in practice

When a question asks whether two events are independent, give a numerical comparison rather than just a sentence. Calculate both sides of one independence condition, then compare them.

Using the multiplication test, for example, calculate P(AB)P(A\cap B) and P(A)P(B)P(A)P(B). Equal values show that the events are independent; unequal values show that they are dependent. Alternatively, calculate P(AB)P(A\mid B) and compare it with P(A)P(A).

Make sure the conclusion matches the calculation. One common trap: mutually exclusive events aren’t automatically independent. If P(A)>0P(A)>0 and P(B)>0P(B)>0, mutually exclusive events have P(AB)=0P(A\cap B)=0 but P(A)P(B)>0P(A)P(B)>0, so they fail the independence test.

Two-way table for independent events AA and BB

EventBBBB'Total
AA0.200.200.40
AA'0.300.300.60
Total0.500.501.00

Probability appears across many disciplines, but its interpretation depends on the context. A calculation may show dependence between two medical or social categories, yet it cannot explain the cause by itself. Here, the distinction between mathematical evidence and subject-specific reasoning is useful rather than artificial.

4.13

BAYES' THEOREM FOR UP TO THREE EVENTS

HL

Reversing a conditional probability

Bayes' theorem combines prior and conditional probabilities to reverse the direction of conditioning. Put simply, it answers questions such as: if the test is positive, how likely is each possible cause?

Begin with two mutually exhaustive alternatives, BB and BB'. Mutually exhaustive events have a union that covers the whole sample space, so one of them must occur. For two alternatives, Bayes' theorem is

P(BA)=P(B)P(AB)P(B)P(AB)+P(B)P(AB)P(B\mid A)=\frac{P(B)P(A\mid B)}{P(B)P(A\mid B)+P(B')P(A\mid B')}

The denominator gives the total probability of AA. It splits this probability into the routes through BB and BB'. Students often miss this step: Bayes is simply conditional probability combined with a complete list of routes to the evidence.

Image

Bayes' theorem for three events

When there are a maximum of three possible underlying events, label them B1B_1, B2B_2, and B3B_3. Here, BiB_i represents the iith possible underlying event for i=1,2,3i=1,2,3. Together, the events should partition the sample space, meaning exactly one occurs.

Bayes' theorem becomes

P(BiA)=P(Bi)P(ABi)P(B1)P(AB1)+P(B2)P(AB2)+P(B3)P(AB3)P(B_i\mid A)=\frac{P(B_i)P(A\mid B_i)}{P(B_1)P(A\mid B_1)+P(B_2)P(A\mid B_2)+P(B_3)P(A\mid B_3)}

A tree diagram usually provides the clearest working method. Place the three possible sources, groups, machines, diagnoses, or categories on the first set of branches, then put the observed event AA and its complement on the second set. Multiply along each path. Finally, divide the path you want by the total of all paths ending in AA.

Image

Link with independence

Bayes' theorem is much less interesting when AA is independent of the underlying event. If AA is independent of BiB_i, then P(ABi)=P(A)P(A\mid B_i)=P(A) for each relevant ii. Observing AA therefore doesn’t update the probabilities of the BiB_i events. This is the SL independence idea in a new setting.

Base rates matter when Bayes' theorem is used in medical screening, economics, quality control, and other real systems. If false positives are common, a rare condition may still be unlikely after a positive test. The mathematics is applicable, but its value depends on whether the model is valid and the probabilities are reliable.

Were those notes helpful?

distributions Distributions

statistics Statistics