A school has students in grade 10, students in grade 11 and students in grade 12. A researcher wants a representative sample of students to investigate weekly study time.
State the most appropriate sampling method for ensuring that each grade is represented proportionally.
Determine the number of students that should be selected from each grade.
Explain why selecting all students from grade 10 would be unsuitable.
0
The times taken by customers to complete an online form are summarized in the table.
Completion time t [min] | Frequency |
|---|---|
0 ≤ t < 10 | 4 |
10 ≤ t < 20 | 7 |
20 ≤ t < 30 | 9 |
30 ≤ t < 40 | 5 |
Total | 25 |
Write down the modal class.
Estimate the mean completion time.
Explain why the value found in part (b) is an estimate.
0
The masses of parcels are displayed in a grouped frequency table. All class intervals have equal width.
Mass interval, m [kg] | Frequency |
|---|---|
0 ≤ m < 5 | 3 |
5 ≤ m < 10 | 8 |
10 ≤ m < 15 | 12 |
15 ≤ m < 20 | 4 |
20 ≤ m < 25 | 3 |
State the modal class.
Calculate the percentage of parcels with mass less than kg.
Explain two features that distinguish a histogram of these data from a bar chart.
0
The delivery times, in minutes, for ten orders are
.
For these data, and .
Determine whether is an outlier.
The value is found to be a recording error and is removed. Calculate the mean delivery time of the remaining orders.
0
The mean temperature of a set of measurements in degrees Celsius is , and the population standard deviation is . Each temperature is converted to degrees Fahrenheit using .
Calculate the mean of the converted temperatures.
Calculate the population variance of the converted temperatures.
State the effect of adding on the standard deviation.
0
The cumulative frequency graph shows the journey times, in minutes, of commuters.

Estimate the lower and upper quartiles.
Estimate the median journey time.
Hence estimate the interquartile range.
Estimate the number of commuters whose journey time is greater than minutes.
0
Box-and-whisker diagrams show the scores of two groups, A and B, on the same test.
Statistic | Group A | Group B |
|---|---|---|
Minimum | 42 | 40 |
Q1 | 48 | 45 |
Median | 54 | 58 |
Q3 | 60 | 65 |
Highest non-outlier | 66 | 76 |
Outlier | None | 92 |
Compare the typical scores of the two groups.
Compare the spread of the middle half of the scores.
Suggest which group is more consistent with a normal distribution. Justify your answer.
0
A company has an ordered list of employees. A researcher uses systematic sampling to select employees. A random starting position of is chosen.
Determine the sampling interval.
Write down the first three and the final employee positions selected.
The list repeats a pattern of job roles every positions. Explain why this may make the sample biased.
0
The table shows the number of faults found in a batch of electronic devices. The frequency corresponding to two faults is . The mean number of faults is .
Number of faults | Frequency |
|---|---|
1 | 3 |
2 | k |
3 | 5 |
4 | 2 |
Find the value of .
Find the population standard deviation of the number of faults.
new variable is formed by adding to every number of faults. State the mean and population standard deviation of the new variable.
0
A data set has five-number summary
.
A new data set is formed using .
Determine the five-number summary of .
Find the interquartile range of .
An additional value is transformed to a value of . Determine whether its transformed value is an outlier of .
0
A data set has lower quartile and upper quartile , where . The data value is being tested as a possible upper outlier.
When , determine whether is an outlier.
Find the value of for which lies exactly on the upper outlier fence.
State whether is classified as an outlier when .
0
A grouped frequency table summarizes the durations of telephone calls. The frequency in the interval is . The estimated mean duration is minutes.
Call duration t [min] | Frequency |
|---|---|
0 ≤ t < 10 | 2 |
10 ≤ t < 20 | 5 |
20 ≤ t < 30 | p |
30 ≤ t < 40 | 3 |
Find .
State the modal class.
Give two reasons why the calculated mean may differ from the mean of the original call durations.
0
The cumulative frequency graph shows the waiting times, in minutes, of patients at a clinic.

Estimate the lower quartile, median and upper quartile.
Using a minimum of minutes and a maximum of minutes, draw a box-and-whisker diagram for the waiting times.
The th percentile is approximately minutes. Estimate the number of patients who waited more than minutes.
0
A university has science students, arts students and business students. Researchers want a sample of students. They set proportional targets for each faculty, email all students and accept the first volunteers until each target is filled.
Determine the proportional target for each faculty.
State the sampling method actually used and explain why it is not proportional stratified random sampling.
Explain why increasing the sample size would not necessarily remove the bias.
Suggest one change that would make the sample less biased while retaining proportional representation.
0
Group A contains observations with mean and population standard deviation . Group B contains observations with mean and population standard deviation . The groups are combined.
Calculate the mean of the combined group.
Calculate the population standard deviation of the combined group.
Explain why the combined standard deviation is greater than the standard deviation of either group.
0
A population data set contains values. Its recorded mean is and its recorded population variance is . One value was recorded as but should have been .
Calculate the corrected mean.
Calculate the corrected population variance.
Calculate the corrected population standard deviation.
Explain why correcting the value reduces the standard deviation.
0
A conservation organization wishes to estimate the number of hours spent each month by volunteers in three nature reserves. The reserves have , and registered volunteers respectively. A proportional stratified random sample of volunteers is selected.
State the population in this investigation.
State whether the number of hours volunteered is discrete or continuous.
Determine the number of volunteers that should be selected from each reserve.
The organization has a complete numbered list of the volunteers in each reserve. Explain how the required volunteers could be selected so that the sample is proportional stratified and random.
coordinator instead surveys the first volunteers arriving at a weekend training event. Explain why increasing this convenience sample to volunteers would not necessarily remove the bias.
0
The cycling times, minutes, of participants are summarized below.
| Time interval | |||||
|---|---|---|---|---|---|
| Frequency |
Write down the modal class.
Estimate the mean cycling time.
Calculate the percentage of participants whose cycling time was at least minutes.
model estimates the energy used by a participant as , where is measured in kilojoules. Hence estimate the mean energy used.
Explain why the mean found in part (a)(ii) is only an estimate.
0
A sensor recorded the daily number of litres of water used by a greenhouse as
For these data, and . The final value is later found to have been entered incorrectly.
Calculate the interquartile range.
Determine the lower and upper outlier fences.
Determine whether the recorded value is an outlier. Justify your answer.
The value should have been . Calculate the corrected mean water use.
Before the recording error was discovered, explain why the median would have been a more representative measure of typical water use than the mean.
0
The cumulative frequency graph shows the operating lifetimes, in hours, of rechargeable batteries.

Use the graph to estimate each of the following:
lower quartile;
median;
upper quartile.
Hence estimate the interquartile range.
The graph indicates that batteries lasted at most hours. Determine the number of batteries that lasted more than hours.
State the th percentile and interpret its meaning in context.
0
The box-and-whisker diagrams compare assessment scores for classes A and B. Class A has five-number summary . Class B has non-outlying five-number summary , with outliers at and .

For the two classes,
compare the median scores;
calculate and compare the interquartile ranges.
Based only on the interquartile ranges, state which class has more consistent scores. Give a reason.
Suggest which class is more consistent with a normal distribution. Justify your answer using two features of the box plots.
At least what percentage of the non-outlying scores in class B are greater than or equal to ? Explain your answer.
0
A set of salinity measurements, , has mean , population standard deviation , lower quartile , median and upper quartile . A calibrated sensor reports
Calculate the
mean of the calibrated measurements;
population standard deviation of the calibrated measurements.
Calculate the population variance of the calibrated measurements.
Determine the lower quartile, median and upper quartile of the calibrated measurements.
Explain separately how multiplying by and subtracting affect the spread of the data.
0
The increases in height, cm, of seedlings are summarized below.
| Height increase | |||||
|---|---|---|---|---|---|
| Frequency |
Find the value of .
Write down the modal class.
Estimate the mean increase in height.
Calculate the percentage of seedlings with an increase of less than cm.
Describe two features that a correctly drawn histogram of these data must have.
0
The numbers of minutes spent resolving technical-support cases are
.
For these data, the calculator quartiles are and .
Calculate the interquartile range and the two outlier fences.
Hence classify any outliers.
Calculate the change in the mean if the -minute case is removed.
Assess whether identifying the -minute value as an outlier is sufficient justification for removing it from the investigation.
0
At a city sustainability festival, researchers ask the first visitors to a recycling station to report their weekly household waste. Of these visitors, submit a form. Twelve respondents leave the waste question blank. The missing entries are recorded as zero, giving a reported mean of for all forms.
State the intended population if the researchers wish to draw conclusions about all households in the city.
State the sampling method used to approach the visitors.
Explain two ways in which the obtained responses may be biased.
The researchers decide to exclude the missing entries rather than treat them as zero. Calculate the mean for the recorded waste values.
Explain why consistently treating every missing entry as zero could make the procedure reliable but not valid.
0
Two machines produce metal rods. Machine A produces rods with mean length cm and population standard deviation cm. Machine B produces rods with mean length cm and population standard deviation cm. The rods are treated as one population.
Calculate the sum of the lengths of all rods.
Hence calculate the combined mean length.
Calculate the population standard deviation of the lengths of all rods.
Every rod length is converted using . Determine the mean and population standard deviation of .
0
A population data set contains measurements with mean . Its lower and upper quartiles are and respectively. One additional measurement, , is added to the data set. The mean of the enlarged data set is .
The mean of the measurements is .
Form an equation in and hence find .
Using the quartiles of the original data set, determine whether is an outlier.
Find the greatest possible mean of the enlarged data set if the added value must not exceed the original upper outlier fence.
Explain why a complete outlier analysis of the enlarged data set may require its quartiles to be recalculated.
0
A laboratory records the operating lives of prototype batteries. The recorded mean is hours and the recorded population standard deviation is hours. During a review, one recorded value of hours is found to be a transcription error; the correct value is hours. The prototypes were supplied by engineers who volunteered to participate in the study.
Calculate the corrected mean operating life.
Hence calculate the corrected population standard deviation.
Evaluate the claim that increasing the number of voluntarily supplied prototypes would, by itself, make the results representative of all batteries produced by the laboratory.
0
The masses, in kilograms, of discarded material collected from households are summarized in a grouped frequency table. The class intervals are , , and , with respective frequencies , , and . The estimated mean mass is kg.
Mass [kg] | Frequency |
|---|---|
4 | |
10 | |
6 |
Determine the value of .
State the modal class.
Estimate the population standard deviation of the masses.
disposal score is defined by . Find the estimated mean and population standard deviation of the scores. If you did not obtain an answer to part (b)(i), use kg.
Explain why the mean mass obtained from the table is an estimate.
0
A hospital has an ordered list of staff members. A researcher plans to select staff members using a systematic sample. The list consists of blocks of roster positions. In each block, one specified position is occupied by a night-shift worker; assignments at all other positions may vary between blocks. The hospital employs day-shift staff, evening-shift staff and night-shift staff.
Determine the systematic sampling interval.
Explain how this roster arrangement could produce a severely biased sample.
The researcher instead uses a proportional stratified random sample by shift. Determine the number selected from each shift.
Give one advantage and one limitation of the stratified method in this context.
0
A council asks randomly selected residents to rate a new public park on a scale from to . There are responses, with a total score of . A spreadsheet replaces each of the missing responses by and reports a mean score of . The council believes every missing score, if known, would lie between and inclusive.
Calculate the mean score of the residents who responded.
Explain why replacing a missing response by zero is not valid.
Determine the lowest and highest possible mean score for all selected residents under the council's belief.
Explain why the original random selection does not guarantee that the responses are unbiased.
0
Two machines fill bottles with juice. Their five-number summaries, in millilitres, are shown below.
Machine A:
Machine B:
A separate test bottle contains ml.
Machine | Minimum / mL | Lower quartile / mL | Median / mL | Upper quartile / mL | Maximum / mL |
|---|---|---|---|---|---|
A | 48 | 50 | 52 | 55 | 59 |
B | 45 | 51 | 54 | 57 | 60 |
Compare the typical fill volumes and the spreads of the middle halves.
Identify which machine has the greater overall range.
Using each machine's stated quartiles, determine how the ml bottle would be classified.
Machine A's readings are recalibrated using . Find the transformed median and interquartile range.
0
Two statistical software packages use different quartile conventions for the same short data set. Package P reports and . Package Q reports and . The data set contains an observation of .
Using Package P, determine whether is an outlier.
Using Package Q, determine whether is an outlier.
Explain why the two classifications do not necessarily indicate that either package is mathematically incorrect.
Suggest how a published report should present this result to avoid misleading readers.
0
A measuring instrument is tested on the same eight reference components on two occasions. On the first occasion, the readings have mean units and population standard deviation units. Every second-occasion reading is exactly units greater than its corresponding first-occasion reading. The certified value for each component is close to units.
Find the mean and population standard deviation of the second-occasion readings.
State the change in the population variance.
Discuss separately what the results suggest about the reliability and validity of the instrument.
technician corrects each second reading using . Find the corrected mean and standard deviation, and state one limitation of concluding that the instrument is now valid.
0
An international organization compares monthly household energy costs. In Country A, the mean cost is local currency units with population standard deviation . The conversion to a common unit is . In Country B, the reported mean is common units with population standard deviation common units. Country A records taxes in the cost, while Country B excludes taxes.
Convert Country A's mean and population standard deviation to the common unit.
Compare the typical costs and their dispersion using the converted statistics.
Evaluate whether the converted statistics support a valid comparison of the countries' monthly household energy costs. Give a justified conclusion, referring to the converted statistics and the different cost definitions.
Suggest any two distinct pieces of methodological information that should be checked before the organization publishes a comparison.
0
The number of defects found on a component takes the values , , , and . Their respective frequencies are , , , and . The mean number of defects is . The quality score is defined by , where is the number of defects.
Number of defects | Frequency |
|---|---|
0 | 3 |
1 | 5 |
2 | k |
3 | 4 |
4 | 2 |
Determine the value of .
State the mode.
Calculate the population standard deviation of the number of defects.
quality score is defined by , where is the number of defects.
Find the mean and population standard deviation of the quality score. If you did not obtain an answer to part (b), use .
Explain why the most frequent defect count produces the most frequent quality score despite the negative multiplier.
0
The number of faults, , found in each of manufactured panels is summarized below.
| Frequency |
The mean number of faults is .
Form two simultaneous equations in and .
Hence find and .
State the mode of the number of faults.
Calculate the population variance and population standard deviation of the number of faults.
quality score is defined by . Determine the mean and population standard deviation of the quality scores.
0
A box plot for a data set has non-outlying five-number summary
Two additional values, and , are plotted separately. A transformed data set is defined by .

Calculate the interquartile range and the outlier fences for .
Verify that both additional values are outliers.
Determine the transformed non-outlying five-number summary, in increasing order.
Transform the two outliers and verify using the transformed outlier fences that they remain outliers.
Explain why any non-zero linear transformation preserves whether a value lies beyond an outlier fence.
0
A public transport authority employs staff in four districts. The district populations are , , and , giving employees in total. A proportional stratified random sample of employees is required.
Calculate the unrounded proportional allocation for each district.
Use the largest-remainder method to obtain integer allocations with total .
Explain why the employees must be selected randomly within each district.
As an alternative, the authority considers a systematic sample of employees from the complete ordered list. Determine the sampling interval and, for a random starting position of , the first three and final selected positions.
The ordered list repeats a work-shift pattern every positions. Explain why this threatens the systematic sample.
0
The daily masses, kg, of material processed by laboratory units are grouped as follows.
| Mass interval / kg | |||||
|---|---|---|---|---|---|
| Frequency |
Using the grouped data, estimate the
mean mass;
population standard deviation.
Every observation is within kg of its class midpoint. Determine the smallest interval that is guaranteed to contain the exact mean of the ungrouped data.
calibration error is discovered: every mass should be increased by kg. State the corrected estimated mean and population standard deviation.
State one limitation of using the grouped standard deviation to describe the original measurements.
0
The table gives cumulative frequencies for the completion times, minutes, of computer tasks.
| Cumulative frequency |
Assume linear interpolation between consecutive points.

Use linear interpolation to estimate the
lower quartile;
median;
upper quartile.
Hence estimate the interquartile range.
Use these estimates to determine the outlier fences. Hence explain whether any task time in the displayed range could be classified as an outlier.
Estimate the percentile rank of a completion time of minutes.
State the assumption about the observations within each interval that is made when using linear interpolation.
0
A data set has lower quartile and upper quartile , where . An additional observation has value . Let be the data set obtained by transforming every value in using . The additional observation is transformed separately in part (b)(ii).
Find the upper outlier fence for .
Hence show that is an outlier of .
Find , and the interquartile range of .
Determine whether the transformed additional observation is an outlier relative to .
Explain why a transformation , where , preserves whether an observation is an outlier.
0
A transport authority combines journey-speed data from two regions. Region A has observations with mean km h and population standard deviation km h. Region B has observations with mean km h and population standard deviation km h.
Calculate the mean speed of the combined data set.
Calculate the population standard deviation of the combined data set.
All Region B speeds are recalibrated by subtracting km h. Calculate the population standard deviation after the recalibrated data are combined with Region A.
Explain why the recalibrated combined standard deviation is smaller.
0
The water used by apartments in one day is grouped into the intervals , , and , measured in cubic metres. The respective frequencies are , , and . No exact individual values are available.
Class interval for water use [m³] | Frequency |
|---|---|
8 | |
14 | |
12 | |
6 |
Estimate the mean daily water use.
State the modal class.
Show that the actual mean must be at least cubic metres and less than cubic metres.
Hence state the greatest possible absolute error in the midpoint estimate and explain when it is attained, and why an error of cubic metres on the upper side cannot be attained.
0
A population data set contains observations with mean and population variance . Eight of the observations have sum and sum of squares . The remaining observations are and , where .
Show that and .
Hence determine and exactly.
transformed data set is defined by . Find its mean and population variance.
Explain why interchanging the labels and would not change any descriptive statistic of the data set.
0
A conservation organization has registered volunteers in three regions: coastal, forest and mountain volunteers. It requires a proportional stratified sample of volunteers for a survey.
Region | Registered volunteers / volunteers | Unrounded proportional allocation / volunteers (student calculates) | Initial allocation / volunteers (student calculates) | Final allocation / volunteers (student calculates) |
|---|---|---|---|---|
Coastal | 137 | |||
Forest | 89 | |||
Mountain | 64 | |||
Total | 290 |
Calculate the unrounded proportional allocation for each region.
Determine integer allocations that total , using the largest-remainder method.
Within each region, the organization accepts the first volunteers who reply to an email until the allocation is filled. Explain why this is not a stratified random sample.
State one modification that would retain the regional allocation while reducing selection bias.
0
A company has production employees, each with an annual salary of monetary units, and executives, each with an annual salary of monetary units. A public report states only that the company's mean salary is high and uses this to describe the pay of a typical employee.
Calculate the mean and median salaries.
Calculate the population standard deviation of the salaries.
Using the quartile convention that separates the lower and upper halves of the ordered data, determine the quartiles and classify the executive salaries using the outlier rule.
Evaluate whether the report's use of the mean is an appropriate description of a typical employee's salary and whether the executive salaries should be deleted as outliers.
0
Two sensors measure the same type of industrial pressure. Sensor A provides readings with mean kPa and population standard deviation kPa. Sensor B provides readings with mean kPa and population standard deviation kPa. A calibration investigation shows that kPa must be added to every reading from sensor B.
For the corrected readings from sensor B, calculate the following:
mean;
population standard deviation.
The corrected readings from both sensors are combined. Calculate the combined mean and population standard deviation.
Before calibration, the combined mean was kPa. Calculate the population standard deviation of all uncorrected readings.
Explain why correcting sensor B reduces the combined standard deviation even though neither sensor's individual standard deviation changes.
0