Clastify logo
Clastify logo
Exam prep
Exemplars
Review
HOT

Bivariate Statistics

Master IB Math AA Bivariate Statistics with notes created by examiners and strictly aligned with the syllabus.

IB Syllabus Requirements for Bivariate Statistics

4.4

Linear correlation, scatter diagrams and regression of y on x

4.10

Regression of x on y

4.4

LINEAR CORRELATION, SCATTER DIAGRAMS AND REGRESSION OF Y ON X

Bivariate data and scatter diagrams

Bivariate data is a data set where every observation contains two linked numerical variables. Each observation is usually plotted as a point on a coordinate grid, with one variable on the horizontal axis and the other on the vertical axis.

A scatter diagram is a statistical graph that shows bivariate data as separate plotted points, making any possible relationship between the variables visible. It isn’t just a picture. Before trusting a numerical result, check the shape of the cloud of points.

Image

Linear correlation is a statistical association where two variables tend to vary together in a pattern that a straight line describes reasonably well. Points that slope upwards from left to right show positive correlation; points that slope downwards show negative correlation. If there is no clear straight-line trend, the data has zero or no linear correlation. We describe the strength as weak, moderate or strong, depending on how closely the points cluster around a straight line.

Pearson's product-moment correlation coefficient

Pearson's product-moment correlation coefficient is a dimensionless numerical measure of the direction and strength of a linear association between two numerical variables. In this course, it is usually calculated using technology.

The key range is

1r1-1 \le r \le 1

An rr value close to 11 shows strong positive linear correlation, while a value close to 1-1 shows strong negative linear correlation. Values close to 00 indicate little or no linear correlation. Take care with that last description: an rr value close to 00 means there is no useful straight-line relationship, not that there is necessarily no relationship at all. A curved pattern can produce a small Pearson value.

Use technology to find rr in practice. A hand calculation may help explain what is happening, but it isn’t the efficient approach in an examination. If you need to judge whether a correlation is statistically significant, the question will supply any required critical value of rr. Compare it with the size of r|r|. If the calculated value is large enough relative to the given critical value, treat the linear correlation as significant in that context.

Lines of best fit by eye and the mean point

A mean point has coordinates equal to the two sample means of the variables:

(xˉ,yˉ)(\bar{x},\bar{y})

A line of best fit is a straight line drawn to represent the central trend in a scatter diagram. For this syllabus, a line drawn by eye should pass through the mean point and balance the points reasonably well. Don’t force it through the origin. Starting at the origin only makes sense when the context shows that zero in one variable should genuinely correspond to zero in the other.

Image

Sensible lines drawn by eye may differ slightly, and that is fine. Ignoring the mean point or drawing a line that misses the visible trend is not.

Correlation is not causation

Causation is a relationship where a change in one variable directly produces a change in another variable. Correlation by itself does not establish causation. For example, a positive correlation could result from a third variable, reverse dependence, a shared trend over time or coincidence in a small data set.

Validity matters here. Statistics may reveal an association, but context is needed to interpret it. A scatter diagram and an rr value can suggest a useful model; neither proves the mechanism behind the relationship.

Regression line of yy on xx

A regression line of yy on xx is a straight-line model fitted to bivariate data to predict the value of yy from a given value of xx. For this course, use technology to find its equation.

The guide writes the regression equation as

y=ax+by=ax+b

The parameter aa gives the predicted change in yy when xx increases by one unit. The parameter bb is the intercept on the yy-axis. Only interpret it when x=0x=0 makes sense in context and is reasonably close to the data. An intercept can be technically correct but meaningless in real life.

Prediction, interpolation and extrapolation

Interpolation uses a fitted model to estimate a value inside the range of the observed data. Extrapolation uses a fitted model to estimate a value outside that range. Interpolation is usually more defensible because it applies the model where actual data has informed it.

Image

Extrapolation is risky because the straight-line pattern may not continue. Data might appear linear over a short interval, then level off, curve, reach a natural limit or stop making sense. Here, an approximation may be useful, but it is not guaranteed to be true.

When working with the regression line of yy on xx, use xx to predict yy. Don’t use the same line backwards to predict xx from yy unless the question specifically justifies doing so. Regression is directional: a line fitted for predicting vertical errors is not automatically suitable when the prediction is reversed.

4.10

REGRESSION OF X ON Y

Why there is a second regression line

The regression line in 4.4 predicts yy from xx. But some questions reverse the direction and ask you to predict xx from yy. In that case, use the regression line of xx on yy. Don’t simply rearrange the equation for yy on xx.

A regression line of xx on yy is a straight-line model fitted to bivariate data that predicts the value of xx from a given value of yy. One convenient form is

x=cy+dx=cy+d

Use technology to find this equation. The two regression lines answer different prediction questions: the line of yy on xx minimises vertical prediction errors in yy, whereas the line of xx on yy minimises horizontal prediction errors in xx. Unless the correlation is perfect, the lines aren’t the same.

Image

Prediction using the xx on yy line

Use the xx on yy regression equation when you’re given a value of yy and need to estimate xx. Substitute the given value to obtain the predicted value of xx.

Be careful not to run the model backwards. An xx on yy line cannot always predict yy reliably from a value of xx. If the question asks you to predict yy from xx, use the yy on xx regression line instead.

As with any regression model, check whether the prediction involves interpolation or extrapolation. A line may be mathematically neat but still describe the future poorly. This raises a useful TOK question: mathematics allows us to project from present data, but the certainty belongs to the model and not necessarily to the world being modelled.

Were those notes helpful?

distributions Distributions