For Integrated Programme students: Your current school materials, teacher instructions, and assessment scope take precedence because IP topic sequence and depth vary by school. This is an Eclat IP guide, not the O-Level / SEC G3 exam-track guide.
How this chapter applies
Eclat core: data collection, statistical displays, misleading representations, averages, grouped data, quartiles, percentiles, range, IQR, standard deviation, and comparison of data sets form the main route.
School-sensitive extension: interpolation conventions, PMCC, regression, and calculator-specific statistics workflows may vary by school.
2027 national comparison: K310 Topic S1 includes tables and diagrams, cumulative frequency and box plots, grouped mean, quartiles, percentiles, standard deviation for grouped and ungrouped data, and comparison using mean and standard deviation.
Check your school: confirm percentile conventions, calculator model steps, and whether correlation coefficients or regression are assessed.
Q: What does IP EMaths Notes (Upper Sec, Year 3-4): 15) Statistics and Data Handling cover? A: Summarise data with mean, median, mode, and cumulative frequency, and interpret scatter plots with correlation language.
The core idea is simple: Organise the data first; the calculation is easier after the table is clean.
Use it as a working check: Use mean, median, mode, range, IQR, and standard deviation to describe different parts of a dataset. For scatter plots, say correlation, not cause.
Then go one layer deeper: Work through the grouped-table examples to practise cumulative frequency, percentiles, and clear interpretation sentences that compare centre, spread, and pattern.
Statistics questions reward tidy tables and well-labelled graphs. Document your calculator steps so you can replicate them under exam conditions.
Keep the full topic roadmap handy via our IP Maths tuition hub so you can jump into related drills, quizzes, or diagnostics as you move through these notes.
K310 can test tables, bar graphs, pictograms, line graphs, pie charts, dot diagrams, equal-width histograms, stem-and-leaf diagrams, cumulative-frequency diagrams, and box plots. Choose the display that matches the data and task.
Bars should not imply continuous intervals unless the data is grouped.
Dot diagram or stem-and-leaf diagram
Showing individual values while retaining distribution shape.
Include a key for a stem-and-leaf diagram.
Equal-width histogram
Showing grouped continuous data.
Adjacent bars touch; do not treat it as a category bar chart.
Cumulative-frequency diagram
Reading medians, quartiles, percentiles, and counts below a boundary.
Read cumulative totals, not class frequencies.
Box plot
Comparing median, spread, and possible skew between data sets.
It does not show every individual value.
A diagram can mislead if axes are truncated, intervals are unequal but drawn equally, pictograms use area inconsistently, or three-dimensional decoration exaggerates differences. State the design choice and the false impression it creates.
Mean of grouped data: xˉ=∑f∑fx, where f denotes the frequency for each class.
Median for grouped data: locate n/2 on the cumulative frequency column and interpolate within that class.
Quartiles and percentiles: use cumulative frequency or ordered lists depending on the dataset size.
Variance/standard deviation (for raw data): with calculator support, record the key inputs in case you need to justify the button presses.
Correlation language: “positive”, “negative”, “no correlation”; do not say “strong cause”.
Box-and-whisker charts compare spread and medians; comment on both when writing conclusions.
Building organised tables
List raw data in ascending order; tally if the dataset is long.
Build frequency and cumulative frequency columns side by side.
For grouped data, add a midpoint column and a fx column.
For two-variable data (x,y), include x2, y2, and xy columns if you need regression or product-moment correlation coefficient (PMCC) later.
Example table layout
Class
Frequency
Cumulative frequency
Midpoint
fx
0-20
6
6
10
60
20-40
12
18
30
360
40-60
10
28
50
500
60-80
7
35
70
490
80-100
5
40
90
450
Total
40
1860
Grouped mean midpoint checkpoint
When data is grouped into class intervals, the original raw values are no longer visible. The midpoint method gives an estimate, not an exact mean.
Step
What to write
Why it matters
Common trap
Find each midpoint
Add the lower and upper class limits, then divide by 2.
The midpoint stands in for every value in that class.
Using the upper class limit as if every value is at the top of the class.
Multiply by frequency
Calculate fx for each class.
A class with more values should contribute more to the estimated total.
Adding the midpoints without weighting by frequency.
Add the estimated totals
Find ∑fx.
This estimates the total of all raw values.
Treating ∑f and ∑fx as the same type of total.
Divide by total frequency
Use xˉ=∑f∑fx.
The denominator is the number of observations.
Dividing by the number of classes instead of total frequency.
Worked check: in the table above, ∑fx=1860 and ∑f=40, so the estimated mean is xˉ=1860/40=46.5 minutes.
Misconception check: a grouped-data mean is only as accurate as the class grouping allows. If the raw values are available, use the raw values instead of replacing them with midpoints.
Measures of central tendency
Mean: use the ∑fx column. For raw data, type all values into the calculator's statistics mode; state “1-VAR” on the calculator if required.
Median: for raw data, pick the middle value(s). For grouped data, locate n/2 within the cumulative frequencies and interpolate. For the example above, n=40, so n/2=20 falls in the 40−60 class: median≈44 minutes (linear interpolation).
Mode: from grouped data, quote the modal class. If the mode is needed more precisely, use the mode estimation formula (optional at IP level).
Measures of spread
For grouped data, use each class midpoint as the representative value. Calculate the estimated mean and mean square from the frequency table, then take the square root of mean square minus the square of the mean. Because the original values are unknown, grouped-data standard deviation is an estimate.
When comparing two data sets, compare both centre and spread. The larger mean indicates the higher average, while the smaller standard deviation indicates greater consistency around that mean.
Range: highest minus lowest.
Interquartile range (IQR): Q3−Q1; less sensitive to outliers than the range.
Variance / standard deviation: for ungrouped data, use calculator output (write down the xˉ and σ2). For grouped data, use the midpoints with the frequency column.
Measure-choice checkpoint
Before writing a comparison sentence, decide whether the question is asking about centre, spread, or unusual values. Pick the statistic that matches that purpose.
Question focus
Better statistic
Why it fits
Common trap
Typical value with no obvious outlier
Mean or median
Mean uses every value; median gives the middle position.
Quoting all three averages without saying what they show.
Typical value when the data is skewed or has an outlier
Median
Median is less affected by one very large or very small value.
Letting one outlier pull the mean and calling it typical.
Consistency or spread of the middle half
IQR
It compares the central 50% of the data.
Using range when the question asks about consistency.
Overall spread including extremes
Range
It uses the smallest and largest values.
Treating range as typical spread when the extremes are unusual.
Calculator comparison of raw datasets
Standard deviation
Smaller standard deviation means values are more tightly clustered around the mean.
Saying lower standard deviation means lower scores; it means less spread.
Worked example: if two classes have similar medians but Class A has IQR 12 and Class B has IQR 28, say their typical values are similar, but Class A is more consistent because the middle half of its data is less spread out.
Worked example - Cumulative frequency percentile
Using the grouped data table above (total n=40), estimate the 90th percentile.
Locate the 90th percentile position: 0.90n=36. The 36th value lies in the 80−100 class because the cumulative frequency is only 35 by the end of the 60−80 class, then reaches 40 by the end of the 80−100 class.
Work inside that class. The lower class boundary is 80, class width is 20, frequency =5, and the cumulative frequency before the class is 35.
So the 90th percentile is approximately 84 minutes.
Cumulative frequency class checkpoint
To choose the class for a percentile, scan down the cumulative-frequency column and stop at the first cumulative frequency that is at least the target position. For the 36th value, 35 is still too small, so you move to the next class.
Common trap: do not choose the class just before the cumulative frequency crosses the target. That previous class contains values up to the 35th observation only.
Grouped-data interpolation checkpoint
After choosing the correct class, write the four interpolation ingredients before substituting. This prevents you from mixing the class frequency with the cumulative frequency.
Ingredient
What it means
For the 90th percentile example
Lower boundary L
Start of the class that contains the target position
80
Previous cumulative frequency c
Number of values before that class
35
Class frequency f
Number of values inside that class
5
Class width w
Width of the class interval
20
Then use the fraction of the way through the class:
estimate=L+ftarget position−c×w.
Worked check: for the 36th value in the 80−100 class, the position is 36−35=1 value into a class containing 5 values. So the estimate is 80+51×20=84.
Misconception check: the denominator is the frequency inside the selected class, not the total sample size. The total sample size was used earlier to find the target position.
Worked example - grouped data (mean, median, quartiles)
Using the table above:
Mean = 401860=46.5 minutes.
Median: cumulative frequencies 6,18,28,35,40. The 20th observation sits in the 40−60 class; interpolating gives 44 minutes to the nearest whole minute.
Quartiles: Q1 is the 10th observation (≈32 minutes), Q3 is the 30th observation (≈61 minutes
Worked example - comparing box plots
Two classes recorded the time spent on revision (minutes per day). Box plots show:
Class A: median 48, IQR 22.
Class B: median 42, IQR 40.
Interpretation:
Class A has a higher median, so its typical student spends more time revising.
Class B has a much wider IQR, so revision times vary more; some students revise significantly more or less than the typical value.
Always comment on both location (median) and spread (IQR, range) when comparing box plots.
Worked example - scatter plot and correlation
A set of ten students recorded their hours of supervised revision (x) and mock exam scores (y). The calculator outputs warning that r=0.86.
State: There is a strong positive correlation between hours and mock scores.
Interpretation: More supervised revision tends to align with higher mock scores; however, correlation does not imply that supervision alone causes the improvement. Other factors (student motivation, resources) may contribute.
Use the line of best fit to make predictions only within the observed data range (interpolation). Extrapolation outside the range is risky.
Scatter-plot prediction checkpoint
Before using a line of best fit, decide whether the question asks for description, prediction, or explanation. Each one needs different wording.
Task wording
What to write first
When to use the line
Common trap
"Describe the correlation"
Direction and strength, such as strong positive correlation
Do not calculate a prediction unless asked.
Writing a causal sentence instead of a pattern sentence.
"Estimate the value when x = ..."
Check that the x-value lies inside the observed range
Use interpolation from the line of best fit.
Reading from a point on the scatter plot instead of the trend line.
"Predict beyond the data"
State that the estimate is unreliable or risky
Use only if the question insists, and flag extrapolation.
Treating the straight-line trend as guaranteed outside the data range.
"Explain why y increases"
Mention correlation, then say another factor could be involved
The graph alone cannot prove a cause.
Saying x causes y just because the points slope upward.
Worked check: if the plotted hours range from 2 to 10, estimating the score at 6 hours is interpolation and is reasonable. Estimating at 14 hours is extrapolation, so the answer should warn that the trend may not continue.
Misconception check: a high PMCC tells you the points are close to a straight-line pattern. It does not prove that changing x directly causes the change in y.
Practice Quiz
Revisit grouped-data averages, spread measures, and data interpretation skills with a self-marking checkpoint.
Try this
The heights (cm) of 50 students are grouped in classes of width 5cm. Outline the steps you would use to estimate the mean height and the median height.
A box plot for Class C shows median 55, lower quartile 40, upper quartile 75, minimum 25, maximum 90. Summarise what this tells you about the distribution of time spent on co-curricular activities per week.
Calculator output for paired data gives xˉ=6.2, yˉ=72, σx=1.4