IB Syllabus Requirements for Bivariate Statistics
4.4
Linear correlation, scatter diagrams and regression of y on x
4.10
Regression of x on y
4.4
LINEAR CORRELATION, SCATTER DIAGRAMS AND REGRESSION OF Y ON X
Bivariate data is a data set where every observation contains two linked numerical variables. Each observation is usually plotted as a point on a coordinate grid, with one variable on the horizontal axis and the other on the vertical axis.
A scatter diagram is a statistical graph that shows bivariate data as separate plotted points, making any possible relationship between the variables visible. It isn’t just a picture. Before trusting a numerical result, check the shape of the cloud of points.

Linear correlation is a statistical association where two variables tend to vary together in a pattern that a straight line describes reasonably well. Points that slope upwards from left to right show positive correlation; points that slope downwards show negative correlation. If there is no clear straight-line trend, the data has zero or no linear correlation. We describe the strength as weak, moderate or strong, depending on how closely the points cluster around a straight line.
Pearson's product-moment correlation coefficient is a dimensionless numerical measure of the direction and strength of a linear association between two numerical variables. In this course, it is usually calculated using technology.
The key range is
An value close to shows strong positive linear correlation, while a value close to shows strong negative linear correlation. Values close to indicate little or no linear correlation. Take care with that last description: an value close to means there is no useful straight-line relationship, not that there is necessarily no relationship at all. A curved pattern can produce a small Pearson value.
Use technology to find in practice. A hand calculation may help explain what is happening, but it isn’t the efficient approach in an examination. If you need to judge whether a correlation is statistically significant, the question will supply any required critical value of . Compare it with the size of . If the calculated value is large enough relative to the given critical value, treat the linear correlation as significant in that context.
A mean point has coordinates equal to the two sample means of the variables:
A line of best fit is a straight line drawn to represent the central trend in a scatter diagram. For this syllabus, a line drawn by eye should pass through the mean point and balance the points reasonably well. Don’t force it through the origin. Starting at the origin only makes sense when the context shows that zero in one variable should genuinely correspond to zero in the other.

Sensible lines drawn by eye may differ slightly, and that is fine. Ignoring the mean point or drawing a line that misses the visible trend is not.
Causation is a relationship where a change in one variable directly produces a change in another variable. Correlation by itself does not establish causation. For example, a positive correlation could result from a third variable, reverse dependence, a shared trend over time or coincidence in a small data set.
Validity matters here. Statistics may reveal an association, but context is needed to interpret it. A scatter diagram and an value can suggest a useful model; neither proves the mechanism behind the relationship.
A regression line of on is a straight-line model fitted to bivariate data to predict the value of from a given value of . For this course, use technology to find its equation.
The guide writes the regression equation as
The parameter gives the predicted change in when increases by one unit. The parameter is the intercept on the -axis. Only interpret it when makes sense in context and is reasonably close to the data. An intercept can be technically correct but meaningless in real life.
Interpolation uses a fitted model to estimate a value inside the range of the observed data. Extrapolation uses a fitted model to estimate a value outside that range. Interpolation is usually more defensible because it applies the model where actual data has informed it.

Extrapolation is risky because the straight-line pattern may not continue. Data might appear linear over a short interval, then level off, curve, reach a natural limit or stop making sense. Here, an approximation may be useful, but it is not guaranteed to be true.
When working with the regression line of on , use to predict . Don’t use the same line backwards to predict from unless the question specifically justifies doing so. Regression is directional: a line fitted for predicting vertical errors is not automatically suitable when the prediction is reversed.
4.10
REGRESSION OF X ON Y
The regression line in 4.4 predicts from . But some questions reverse the direction and ask you to predict from . In that case, use the regression line of on . Don’t simply rearrange the equation for on .
A regression line of on is a straight-line model fitted to bivariate data that predicts the value of from a given value of . One convenient form is
Use technology to find this equation. The two regression lines answer different prediction questions: the line of on minimises vertical prediction errors in , whereas the line of on minimises horizontal prediction errors in . Unless the correlation is perfect, the lines aren’t the same.

Use the on regression equation when you’re given a value of and need to estimate . Substitute the given value to obtain the predicted value of .
Be careful not to run the model backwards. An on line cannot always predict reliably from a value of . If the question asks you to predict from , use the on regression line instead.
As with any regression model, check whether the prediction involves interpolation or extrapolation. A line may be mathematically neat but still describe the future poorly. This raises a useful TOK question: mathematics allows us to project from present data, but the certainty belongs to the model and not necessarily to the world being modelled.