A school has students. The numbers of students in Years 10, 11 and 12 are , and respectively. A researcher wants to take a stratified random sample of students.
State the population and the sample size.
Find the number of students that should be selected from each year group using proportional allocation.
Suggest one reason why asking the first students who enter the school in the morning may give biased results.
0
The following grouped frequency table shows the time, minutes, taken by students to complete a puzzle.
| Time, minutes | ||||
|---|---|---|---|---|
| Frequency |
Write down the modal class.
Estimate the mean time taken.
State why the value found in part (b) is only an estimate.
0
A data set has mean , standard deviation and median . Each value in the data set is transformed to a new value , where
Find the mean of the transformed data set.
Write down the standard deviation of the transformed data set.
Find the variance of the transformed data set.
Find the median of the transformed data set.
0
A grouped frequency table for the distance, km, travelled by cyclists is shown.
| Distance, km | ||||
|---|---|---|---|---|
| Frequency |
Find the cumulative frequency up to and the cumulative frequency up to .
State the class interval containing the median.
Estimate the mean distance travelled.
0
The five-number summary for the masses, in kg, of some parcels is
where the values are the minimum, lower quartile, median, upper quartile and maximum respectively.
Find the interquartile range.
Find the upper outlier boundary.
Determine whether the maximum value is an outlier. Give a reason.
0
The ordered data set
has mean .
Find the value of .
Write down the median and the range of the data set.
0
A town has households in three districts. Districts A, B and C contain , and households respectively. A survey of households is to be carried out.
systematic sample is taken from a list of all households. Find the sampling interval.
Find the number of households that should be selected from each district in a proportional stratified sample.
The household list is arranged in a repeating pattern by street type, and every th household is then selected after a random start. Explain why this systematic sample may be biased.
0
A school has students. The numbers of students in Years 10, 11 and 12 are , and respectively. A researcher wants to take a sample of students to investigate the time spent on homework.
Find the number of students that should be chosen from each year group if a stratified sample proportional to year-group size is used.
The researcher instead asks the first students who enter the library after school. State the sampling method used and give one reason why this may be biased.
0
The grouped frequency table shows the time, minutes, taken by students to travel to school.
Class interval, [min] | Frequency |
|---|---|
4 | |
11 | |
18 | |
12 | |
5 | |
Total | 50 |
State the modal class.
Use mid-interval values to estimate the mean travel time.
Briefly explain why the answer to part (b) is an estimate.
0
The following are the recorded waiting times, in minutes, for customers at a service desk.
For these data, a GDC gives and .
Find the interquartile range and the upper outlier boundary.
Determine whether minutes is an outlier.
Give one reason why the value should not automatically be removed from the data set.
0
A data set of temperatures in degrees Celsius has mean , median , variance and interquartile range . Each temperature is converted to degrees Fahrenheit using
Find the mean and median of the Fahrenheit temperatures.
Find the standard deviation and variance of the Fahrenheit temperatures.
Find the interquartile range of the Fahrenheit temperatures.
0
Data set contains values and has mean . Data set contains values and has mean . When the two data sets are combined, the mean is . The combined data set has standard deviation . Each combined value is transformed to
Find the value of .
Find the mean of the transformed values .
Find the variance of the transformed values .
0
The following grouped frequency table shows values of a continuous variable .
| Class interval | ||||
|---|---|---|---|---|
| Frequency |
The estimated mean, using mid-interval values, is .
Find the value of .
Write down the modal class.
Find the proportion of the data values which are less than .
0
A data set has mean and standard deviation . A new data set is formed using
where . The data set has mean and standard deviation .
Find the values of and .
Find the variance of .
The median of is . Find the median of .
0
The five-number summaries for the waiting times, in minutes, at two clinics are shown.
Clinic A:
Clinic B:
In each summary, the values are the minimum, lower quartile, median, upper quartile and maximum respectively.
Find the interquartile range for each clinic.
Compare the medians and the spreads of the waiting times at the two clinics.
Use the rule to determine whether the maximum waiting time at Clinic A is an outlier.
0
The cumulative frequency graph shows the marks, , obtained by students in a test.

Estimate the median mark.
Estimate the interquartile range.
Estimate the mark below which of the students scored.
0
The table shows the five-number summaries for the daily number of bicycle rentals from two locations, A and B, over the same period.
Location | Min | Median | Max | ||
|---|---|---|---|---|---|
A | 2 | 8 | 12 | 18 | 30 |
B | 4 | 10 | 14 | 20 | 24 |
Find the range and interquartile range for each location.
Compare the two distributions using the five-number summaries.
0
A university has undergraduate students in four faculties. The numbers of students in Science, Arts, Business and Other faculties are , , and respectively. A stratified random sample of students is required.
Determine the number of students to select from each faculty, using proportional stratified sampling. Use the largest remainder method if needed.
The students are listed alphabetically within each faculty, and a systematic sample is taken by choosing every th student after a random start. This procedure may produce or students. Explain why this may not be as effective as the stratified sample in part (a).
State one feature required for the sample in part (a) to be a stratified random sample.
0
The grouped frequency table shows the masses, kg, of parcels. One frequency is missing. The estimated mean mass, using mid-interval values, is kg.
Mass, m [kg] | Mid-interval value [kg] | Frequency |
|---|---|---|
50 ≤ m < 55 | 52.5 | 6 |
55 ≤ m < 60 | 57.5 | a |
60 ≤ m < 65 | 62.5 | 18 |
65 ≤ m < 70 | 67.5 | 14 |
70 ≤ m < 75 | 72.5 | 4 |
Find the missing frequency .
Using , estimate the standard deviation of the parcel masses.
State the modal class.
0
The cumulative frequency graph shows the delivery times, minutes, for packages. The largest recorded delivery time in the original data set was minutes. A later check finds an additional delivery time of minutes.

Estimate the number of packages delivered in less than minutes.
Estimate the interquartile range of the original data set.
Using your answer to part (b), determine whether the additional delivery time of minutes would be classified as an outlier.
0
The following data give the pH values of water samples.
Use your GDC to find the mean and standard deviation of the pH values.
Each pH value is transformed to . Find the mean and standard deviation of the transformed values.
constant is added to each original pH value so that the new mean is . Find .
0
For a set of measurements from production line B, the coded variable
is used, where is the measurement in millimetres. A GDC gives the mean of as and the variance of as . For production line A, the standard deviation of the measurements is mm.
Find the mean and standard deviation of the original measurements from line B.
technician subtracts mm from every measurement from line B. Find the new mean and standard deviation for line B.
After the technician's adjustment, determine which production line has less variation in measurements.
0
The heights, cm, of seedlings are recorded in the grouped frequency table.
| Height, cm | |||||
|---|---|---|---|---|---|
| Frequency |
The standard deviation of the original heights is given as cm.
For each seedling, define a new variable by .
Find the cumulative frequency up to and the cumulative frequency up to .
Write down the median class and the modal class.
Use mid-interval values to estimate the mean height of the seedlings.
State why the answer to part (b)(i) is only an estimate.
new variable is defined by .
Using your answer to part (b)(i), estimate the mean of . If you did not obtain , use this value.
Find the standard deviation of .
0
A city council wants to survey residents about a new public transport route. The city has adult residents, divided into three districts as shown.
| District | North | Central | South |
|---|---|---|---|
| Number of adult residents |
A sample of residents is to be selected.
State the population for this survey.
Find the number of residents that should be selected from each district in a proportional stratified sample.
The council instead selects every th name from an alphabetical list of all adult residents, after choosing a random starting point. State the sampling method used.
Explain why asking the first people leaving the central train station on a Monday morning may give biased results.
Some residents in the South district do not have internet access and so cannot complete an online version of the survey. State one possible effect on the reliability of the results.
0
The five-number summaries for the daily water usage, in litres, of two households over the same month are shown.
Household A:
Household B:
In each summary, the values are the minimum, lower quartile, median, upper quartile and maximum respectively.
Find the range and interquartile range for Household A.
Find the range and interquartile range for Household B.
Use the rule to determine whether the maximum value for Household A is an outlier.
Compare the typical daily water usage of the two households.
Compare the spread of the daily water usage of the two households.
0
The following ordered data set contains seven values:
The mean of the data set is .
Find the value of .
Write down the median and the range of the data set.
For this question, take and to be the medians of the lower and upper halves of the ordered data after excluding the overall median.
Find the interquartile range of the data set.
new value, , is added to the data set. Determine whether is an outlier compared with the original data set.
0
The times, minutes, taken by a group of students to complete an online task are shown in the grouped frequency table.
| Time, minutes | ||||
|---|---|---|---|---|
| Frequency |
There are students in the group.
Find the value of .
Write down the modal class.
Use for this part.
Estimate the mean time taken to complete the task.
Find the percentage of students who took less than minutes.
Each task time is increased by minutes because of a delay in logging in.
Find the new estimated mean time.
State what happens to the standard deviation of the task times.
0
For this question, take and to be the medians of the lower and upper halves of the ordered data, after excluding the overall median.
The ordered data set is
where and is an integer.
Find the median and the interquartile range in terms of the given data.
Find the least possible value of such that is an outlier.
Explain why changing from to a larger value does not change the median or the interquartile range.
0
A wildlife charity has annual pass holders. The pass holders are divided into three categories: local adult, family and concession. The numbers in these categories are , and respectively. The charity wants to survey pass holders about the time they spend in the park on each visit. A proportional stratified sample is to be selected, with each category represented in the same proportion as in the population.
A pilot survey of pass holders produced the following grouped frequency table for the time, hours, spent in the park on one visit.
| / hours | |||||
|---|---|---|---|---|---|
| Frequency |
The charity decides to use proportional stratified sampling.
Determine the number of pass holders that should be selected from each category.
If a systematic sample of pass holders is taken from the complete list of pass holders, find the sampling interval.
The complete list is ordered by the month in which the pass expires. Suggest one reason why the systematic sample may be biased.
Use the pilot survey data.
Write down the modal class.
Use mid-interval values to estimate the mean time spent in the park.
State why the mean found in part (b)(ii) is only an estimate.
0
The number of hours of sleep, , recorded by athletes on the night before a competition is shown in the grouped frequency table.
| Sleep, / hours | ||||||
|---|---|---|---|---|---|---|
| Frequency |
cumulative frequency graph is to be drawn.
Write down the cumulative frequencies at and .
Estimate the median number of hours of sleep.
Use linear interpolation within each class interval.
Estimate the lower quartile and the upper quartile.
Find the interquartile range.
Estimate the th percentile.
Use the grouped frequency table to estimate the mean number of hours of sleep and comment briefly on what comparison with the estimated median suggests about the distribution.
0
In an escape room, the time, minutes, taken by groups to solve the final clue was recorded. The grouped frequency table is shown.
| Time, minutes | |||||
|---|---|---|---|---|---|
| Frequency |
The total number of groups is .
Find the value of .
Write down the class interval containing the median time.
Use and mid-interval values.
Estimate the mean solving time.
Use your GDC to estimate the standard deviation of the solving times.
After the room is redesigned, the time for each group is modelled by , where is the original solving time in minutes. Calculate the new mean and standard deviation of the solving times, and state how the transformation affects these measures.
0
The table shows the battery life, in hours, of two brands of portable charger tested under the same conditions.
Brand | Minimum / h | Lower quartile / h | Median / h | Upper quartile / h | Maximum / h |
|---|---|---|---|---|---|
Brand A | 6 | 9 | 12 | 15 | 24 |
Brand B | 8 | 10 | 13 | 17 | 20 |
Use the five-number summaries in the table.
Write down the median battery life for each brand.
Find the interquartile range for each brand.
Compare the two distributions, using the five-number summaries in the table.
Use the rule for Brand A.
Find the upper outlier boundary for Brand A.
State whether the maximum battery life for Brand A is an outlier. Give a reason.
0
A laboratory records the reaction times, in seconds, for trials.
Use your GDC to find the mean and median reaction time.
For these data, and . Determine whether seconds is an outlier.
The laboratory suspects that seconds was caused by a faulty sensor. State which measure from part (a), the mean or the median, would be less affected if this value were removed.
0
A city council investigates the weekly number of minutes, , that residents spend using a new public bicycle scheme. The city is divided into four districts.
A total sample of residents is to be selected.
District | Population | Mean from respondents [minutes] |
|---|---|---|
A | 1200 | 14.2 |
B | 1800 | 11.8 |
C | 900 | 18.5 |
D | 1500 | 9.6 |
Find the number of residents that should be selected from each district using proportional stratified sampling.
State one condition needed for the sample in part (a)(i) to be a stratified random sample.
Explain why asking only people who arrive at bicycle stations between and would not be an effective method for this investigation.
Use the district means to estimate the mean weekly time for all residents in the city.
student instead calculates the mean of the four district means. Find this value and explain why it is not the most appropriate estimate for the city mean.
District C has the lowest response rate. Suggest how this missing data could affect the estimate in part (b)(i), given that District C has the largest sample mean.
0
The times, minutes, taken by participants to complete a fitness course have mean , median , standard deviation and interquartile range .
A coded variable is defined by
Find the mean and median of the coded values .
Find the standard deviation and interquartile range of the coded values .
Another group of participants completes the same course. Their coded values have mean and median .
Find the combined mean of the coded values for all participants.
Explain why the combined median cannot be found from the two group medians alone.
0
A company has employees working on three shifts.
| Shift | Morning | Afternoon | Night |
|---|---|---|---|
| Number of employees |
The company wants to take a sample of employees to investigate job satisfaction.
Find the number of employees from each shift that should be selected in a proportional stratified sample.
State one requirement for the sample in part (a)(i) to be a stratified random sample.
The company instead lists all employees in a repeating shift pattern: Morning, Afternoon, Night, Morning, Afternoon, Night, and so on. It then selects every th employee after a random start.
Find the sampling interval that would be used for a systematic sample of employees from employees.
Explain why the systematic sample described may not represent the three shifts fairly.
After the sample is selected, six night-shift employees do not return the questionnaire. Their manager says that these non-respondents are more likely to be dissatisfied.
State whether this missing data should simply be ignored. Give a reason.
Suggest one way to reduce this source of bias.
0
The following data give the maximum sound levels, in decibels, recorded at road junctions during one morning.
A GDC gives and for these data.
Use your GDC for the full data set.
(a)(i) Find the mean sound level.
(a)(ii) Find the standard deviation of the sound levels.
(a)(iii) Write down the median sound level.
Use the rule.
(b)(i) Find the upper outlier boundary.
(b)(ii) Determine whether dB is an outlier. Give a reason.
The value dB is removed after it is found to have been caused by roadworks next to the sensor.
(c)(i) After removing the outlier of dB, find the new mean and standard deviation.
(c)(ii) After removing dB, each of these remaining readings is recalibrated by subtracting dB. Find the mean and standard deviation of these recalibrated readings.
0
A plant nursery records the number of seeds germinating in each of trays. Each tray originally contained seeds.
The numbers of seeds that germinated are
.
Use your GDC for the full data set.
Find the mean and standard deviation of the number of seeds germinating.
Find the median number of seeds germinating.
The GDC gives and . Determine whether is an outlier.
The nursery converts each value to a germination percentage using . Find the mean, standard deviation and median of the germination percentages.
The nursery manager says that the tray with germinated seeds should be deleted because it is an outlier.
Give one possible valid reason why the value should not automatically be deleted.
The first trays on the lowest shelf were used for the sample. State the sampling method and explain why this may be biased.
0
The weekly training distances, km, of runners are summarized in the following table.
| Distance, km | |||||||
|---|---|---|---|---|---|---|---|
| Frequency |
cumulative frequency graph is drawn from the table.
Write down the cumulative frequency at and at .
Estimate the median weekly training distance.
Use linear interpolation within classes.
Estimate the interquartile range.
Estimate the th percentile.
Use mid-interval values to estimate the mean weekly training distance, and then compare it with the median found in part (a)(ii). If you did not obtain a median, use km.
0
A research group records the travel times, minutes, of ferry passengers.
| Time, / min | ||||||
|---|---|---|---|---|---|---|
| Frequency |
Time, [min] | Frequency |
|---|---|
4 | |
9 | |
17 | |
22 | |
13 | |
5 |
State the modal class.
Estimate the mean travel time.
Construct the cumulative frequencies at the upper class boundaries.
Using linear interpolation within the relevant classes, estimate the median and the interquartile range.
later check finds one additional recorded travel time of minutes. Using your estimates from part (b)(ii), determine whether this value would be classified as an outlier. If you did not obtain an interquartile range in part (b)(ii), use minutes.
0
A manufacturer measures the fill volume, ml, of bottles from two machines. For Machine A, the five-number summary is
and for Machine B it is
For Machine A, the mean is ml and the standard deviation is ml. A calibration changes each Machine A value to .
Machine | Min [ml] | Q1 [ml] | Median [ml] | Q3 [ml] | Max [ml] |
|---|---|---|---|---|---|
Machine A | 48 | 51 | 54 | 58 | 67 |
Machine B | 44 | 50 | 53 | 56 | 60 |
Find the median, interquartile range and range for Machine A before calibration.
Determine whether the maximum value for Machine A is an outlier.
Find the mean and standard deviation of the calibrated Machine A values.
Find the median and interquartile range of the calibrated Machine A values.
Using the summaries, compare the calibrated Machine A distribution with Machine B in terms of centre and spread.
0
The lifetimes, hours, of a sample of components are grouped as follows.
| Lifetime, hours | |||||
|---|---|---|---|---|---|
| Frequency |
Using mid-interval values, the estimated mean lifetime is hours.
Lifetime [h] | Frequency |
|---|---|
6 | |
18 | |
12 | |
4 |
Find the value of .
State the modal class.
Explain why the mean found from this table is an estimate.
Using , estimate the standard deviation of the lifetimes.
Find the median class.
component is described as unusually long-lasting if its lifetime is more than two estimated standard deviations above the estimated mean. Using your answers above, determine the smallest integer lifetime, in hours, that would be described as unusually long-lasting. If you did not obtain the standard deviation, use hours.
0
Two farms record the masses, kg, of pumpkins harvested on the same day. The five-number summaries are shown.
Farm | Minimum [kg] | Q1 [kg] | Median [kg] | Q3 [kg] | Maximum [kg] |
|---|---|---|---|---|---|
East | 1.8 | 3.1 | 3.6 | 4.0 | 5.2 |
West | 1.2 | 2.4 | 3.7 | 5.0 | 6.3 |
Find the interquartile range for each farm.
Find the range for each farm.
Determine whether the maximum mass for West farm is an outlier.
Comment on whether the West farm distribution may be approximately symmetric, using the five-number summary.
The farms want pumpkins with consistent masses close to kg. Evaluate which farm appears more suitable, based only on these summaries.
0
A national park wants to estimate the mean number of hours spent by visitors on walking trails. Last year, visitor numbers were recorded by entrance.
| Entrance | North | South | East | West |
|---|---|---|---|---|
| Visitors |
A sample of visitors is required.
Entrance | Visitors |
|---|---|
North | 46000 |
South | 28000 |
East | 16000 |
West | 10000 |
Find the exact proportional sample size for each entrance.
Name the sampling technique used in part (a)(i).
State why a complete list of visitors for each entrance (stratum) would be useful.
Explain why a quota sample with the same numbers as in part (a)(i) may still be biased.
Give one example of a convenience sample in this context.
The park later finds that visitors leaving through the West entrance were less likely to respond because mobile reception was poor there. Discuss the possible effect on the estimate of the mean trail time.
0
A data set has values , mean and variance . A new data set is formed by
Show that the mean of is .
Show that the variance of is .
For a particular data set , , , the median is and the interquartile range is .
Find the mean and variance of . If you did not obtain the results in part (a), use and .
Find the median and interquartile range of .
Explain why the negative multiplier does not make the interquartile range negative.
0
A continuous variable is recorded for observations and grouped as shown.
| Class interval | ||||
|---|---|---|---|---|
| Frequency |
Using mid-interval values, the estimated mean is .
Write down an equation involving and using the total frequency.
Write down a second equation involving and using the estimated mean.
Hence find and .
Use and for this part.
Write down the modal class.
Find the proportion of observations for which .
coded variable is defined by .
Find the estimated mean of .
Given that the estimated standard deviation of is , find the estimated standard deviation of .
0
For this question, take and to be the medians of the lower and upper halves of the ordered data, after excluding the overall median.
Consider the ordered data set
where and is an integer.
Find the median, and .
Find the interquartile range.
Find the upper outlier boundary.
Find the least possible value of such that is an outlier.
Explain why increasing beyond does not change the median or interquartile range.
State which measure, the range or the interquartile range, is more affected by increasing .
0
An ordered data set contains nine integer values:
For this data set, the mean is .
Find the value of .
Write down the median.
For this question, take and to be the medians of the lower and upper halves of the ordered data, after excluding the overall median. Use for this part.
Find , and the interquartile range.
Determine whether is an outlier.
The value is replaced by an integer . The mean of the resulting nine values is .
Find .
State one measure from part (b)(i) that is unchanged when is replaced by .
0
A researcher records the lifetimes, hours, of a sample of batteries from a production batch. The grouped frequency table is shown.
| Lifetime, hours | |||||
|---|---|---|---|---|---|
| Frequency |
A second sample from the same production line contains only batteries tested during the first hour after maintenance.
Write down the modal class.
Use mid-interval values to estimate the mean lifetime.
The researcher uses the grouped table to form cumulative frequencies.
Find the cumulative frequency up to and up to .
State the class interval containing the median.
second sample from the same production line contains only batteries tested during the first hour after maintenance.
Explain why this second sample may be biased as evidence about the whole production batch.
Suggest a more appropriate sampling method and justify your choice.
0
A researcher records the lengths, cm, of fish caught and released from a lake. The grouped frequency table is shown.
| Length, / cm | ||||||
|---|---|---|---|---|---|---|
| Frequency |
The total frequency is .
Find the value of .
Write down the modal class.
Use and mid-interval values.
Estimate the mean length of the fish.
Use your GDC to estimate the standard deviation of the lengths.
Use linear interpolation on the grouped data.
Estimate the median length.
Estimate the interquartile range.
fish of length cm is caught the following day. If you did not obtain an interquartile range in part (c)(ii), use cm. Determine whether the fish is an outlier, giving a reason.
0
A technician records the daily energy output, in kWh, of solar panels of the same model. For these values, the mean is , the standard deviation is , and .
A further recorded value of kWh is later found in the data file.
Use the original values.
Find the upper outlier boundary.
State whether kWh is an outlier. Give a reason.
Assume that kWh is added as a th value.
Find the new mean.
Given that the original standard deviation is the population standard deviation, find the new standard deviation.
The recorded value of kWh is found to be erroneous and is replaced by a corrected value kWh, so that the mean of all values remains kWh.
Find .
Find the standard deviation of all values after this correction.
Explain why the standard deviation in part (c)(ii) is less than the original standard deviation.
0
A city library system has registered users in four age groups. The numbers of users are shown in the table.
| Age group | to | to | to | |
|---|---|---|---|---|
| Number of users |
The library wants a sample of users to estimate the mean number of library visits per month.
proportional stratified sample is planned.
Determine the number of users to select from each age group.
Find the sampling interval if a systematic sample of users is taken from the full list of registered users.
The questionnaires returned from the stratified sample are summarized below.
| Age group | to | to | to | |
|---|---|---|---|---|
| Number returned | ||||
| Mean visits per month |
Calculate the unweighted mean number of visits per month for the users who returned questionnaires.
Calculate a weighted estimate of the mean number of visits per month for all registered users, using the population proportions of the four age groups.
Evaluate possible bias in the data collection.
Explain why non-response may bias the unweighted respondent mean.
Evaluate the risk that this systematic sampling method may be biased. Give one reason.
0
A manufacturer measures the width, mm, of metal washers. The coded variable
is used. A GDC gives the mean of as and the standard deviation of as .
For the original widths, a GDC gives , and maximum .
The machine is adjusted so that each width is transformed to , where is measured in mm.
Part (a): Use the coded variable.
Find the mean width of the washers.
Find the standard deviation and variance of the original widths.
Part (b): The machine is adjusted so that each width is transformed to .
Find the mean and standard deviation of the adjusted widths.
Find the adjusted upper quartile and adjusted interquartile range.
Part (c): Use the rule to consider outliers.
Determine whether the original maximum width mm is an outlier.
Hence determine whether the adjusted value corresponding to mm is an outlier.
0
A drinks company checks bottles packed in crates of . There are crates, giving bottles in total. In every crate, the bottle in position is filled by the same nozzle, which is distinct from the nozzles filling the other positions. A quality controller takes a systematic sample of bottles from the ordered list of all bottles, choosing a random starting position from to and then selecting every th bottle.

Verify that the sampling interval is consistent with a sample of bottles.
State the sampling method used.
Find the probability that the systematic sample consists only of bottles from position in their crates.
If the starting position is not or , how many bottles from position are selected?
Explain why this systematic sample may be biased if the nozzle for position is poorly calibrated.
Suggest a more effective sampling method for this situation and justify your answer.
0
The cumulative frequency graph for the lengths, cm, of seedlings is approximated by joining the following points with straight line segments:

Use the graph approximation to estimate the median length.
Estimate the th percentile.
Use the cumulative frequencies to form a grouped frequency table with class intervals of width cm.
Estimate the mean seedling length from this grouped table.
Seedlings shorter than cm are classified as small. Estimate the number of small seedlings, and state one limitation of this estimate.
0
A data logger records temperature readings, in degrees Celsius. The readings have mean and standard deviation . A further reading is then added to the data set.
If and , find the new mean.
Comment on whether is likely to have a larger effect on the mean or on the median.
State one reason why the reading should not automatically be removed.
Let the original readings have mean . Show that after adding one further reading , the new mean is
Hence, treating as an exact target, find the formal value of for which adding would increase the mean from to , and explain why this is not an allowable integer count. If you did not show the result in part (b), use it here.
Explain why the answer in part (a)(i), rounded to significant figures, is consistent with your result in part (c)(i).
0
A museum records the ages, years, of visitors during one afternoon. The ages are grouped in equal class intervals.
| Age / years | ||||
|---|---|---|---|---|
| Frequency |

Estimate the mean age.
Estimate the standard deviation of the ages.
Determine the class containing the median.
Estimate the median using linear interpolation.
publicity report states that the typical visitor is about years old. Evaluate this statement using your results.
0
A data set of breaking strengths, newtons, for ceramic samples has five-number summary
,
where the values are the minimum, lower quartile, median, upper quartile and maximum respectively.
A new variable is defined by .
Use the five-number summary for .
Find the interquartile range and the upper outlier boundary for .
Determine whether the maximum value is an outlier.
Find the five-number summary for .
Consider a general linear transformation , where .
Justify why the interquartile range of is times the interquartile range of .
Hence state why an outlier remains an outlier under any transformation , where .
0
Two laboratories measure the same chemical concentration, mg l, using different instruments. Laboratory A has readings with mean and standard deviation . Laboratory B has readings with mean and standard deviation . Treat each laboratory data set as a population.
Find the mean of the combined readings.
Find the combined standard deviation.
Laboratory B discovers that each of its readings was mg l too high. Find the corrected mean and standard deviation for Laboratory B.
Find the corrected combined mean.
Explain why subtracting from every Laboratory B reading does not change its standard deviation.
After correction, a single combined mean is reported. Discuss one advantage and one limitation of reporting only this value.
0
A data set consists of ordered values with mean , variance , median , lower quartile and upper quartile . A new data set is formed by , where .
Show that the mean of is .
Show that the standard deviation of is .
Explain why the median of is .
Deduce an expression for the interquartile range of .
set of exam scores has mean , standard deviation , median and interquartile range . A school reports scaled scores using . Use the results above to find the mean, standard deviation, median and interquartile range of the scaled scores. If you did not prove the results above, you may still use them.
0