OCR A-level Mathematics A (H240) ยท 2.01 Statistical Sampling
Mini-Lesson
Statistical sampling
OCR section 2.01 is short but it underpins the whole Statistics paper. You need population versus sample, the sampling frame, the main sampling methods with their advantages and disadvantages, and a clear grasp of bias.
You are also expected to be familiar with OCR's pre-release large data set. Work through each screen, answer the questions as you go and collect ⭐ stars. Press Start when you're ready.
Sampling · language
Population, census and sample
Population โ every member of the group being studied. Not necessarily people: it could be every light bulb from a factory, or every day of rainfall data.
Census โ data collected from the whole population. Completely accurate, but usually expensive, slow, and sometimes impossible.
Sample โ a subset of the population, used to estimate what the whole population is like.
Sampling frame โ a list of every member of the population (e.g. a school register), which you draw the sample from.
Why not always a census? Sometimes testing destroys the item. If you crash-test every car you produce, you have no cars left to sell. Sampling is not a compromise here โ it is the only option.
Bias is when a sample systematically misrepresents the population. A bigger sample does not fix bias โ it just gives you a more precise wrong answer.
Quick check
Census or sample?
?A factory tests the lifetime of its light bulbs by running each one until it fails. Why must they use a sample rather than a census?
Sampling · random methods
Random sampling methods
In a random method, chance decides who is chosen โ so it can be free from human bias.
Simple random sampling โ every member of the population has an equal chance of being chosen. Number the sampling frame, then use random numbers. Fair and unbiased, but needs a complete sampling frame and can be slow for large populations.
Systematic sampling โ order the frame, choose a random start, then take every kth member, where k = N/n. Quick and simple, but can go badly wrong if the list has a repeating pattern with the same period as k.
Stratified sampling โ split the population into strata (groups), then sample each stratum in proportion to its size. Best representation, but you must know the strata sizes in advance.
stratum sample size = (stratum size ÷ population size) × total sample sizesystematic sampling interval: k = N / n
Sampling · stratified
Stratified sampling โ worked
Worked example
A school of 1200 students has Year 11: 480, Year 12: 450, Year 13: 270. A stratified sample of 80 is taken.
Sampling fraction = 80 / 1200 = 1/15.
Year 11: (480/1200) × 80 = 0.4 × 80 = 32
Year 12: (450/1200) × 80 = 0.375 × 80 = 30
Year 13: (270/1200) × 80 = 0.225 × 80 = 18
Check they total the sample: 32 + 30 + 18 = 80. ✓ Always do this โ it catches arithmetic slips instantly.
When numbers do not come out whole, round sensibly and adjust so the parts still add to the required total. Say what you did.
Calculate
Your turn โ stratified sample
1A school of 1200 students (Year 11: 480, Year 12: 450, Year 13: 270) takes a stratified sample of 80. How many students come from Year 12?
students
Hint: (450 ÷ 1200) × 80 = 0.375 × 80.
Calculate
Your turn โ the other stratum
2Same school, same stratified sample of 80. How many students come from Year 13 (270 students)?
students
Hint: (270 ÷ 1200) × 80 = 0.225 × 80. Check: your three answers should add to 80.
Sampling · systematic
Systematic sampling โ worked
Worked example
A population of N = 2400 is to give a sample of n = 60.
Interval k = N/n = 2400 / 60 = 40.
Pick a random start between 1 and 40 โ say 17. Then select members 17, 57, 97, 137, …
The 5th member selected is 17 + 4 × 40 = 177.
The danger: if the list has a hidden periodic pattern matching k, the sample is badly biased. Sampling every 7th day of a rota, for example, could pick only Mondays โ and Mondays may be nothing like the rest of the week.
Calculate
Your turn โ the sampling interval
3A systematic sample of 60 is taken from a population of 2400. Find the sampling interval k.
k =
Hint: k = N/n = 2400 ÷ 60.
Calculate
Your turn โ which member?
4With interval k = 40 and a random start of 17, which member of the list is the 5th one selected?
Hint: the selections are 17, 57, 97, 137, … The 5th is 17 + 4 × 40.
Sampling · non-random
Non-random methods
Opportunity (convenience) sampling โ take whoever is available. Very quick and cheap, but highly likely to be biased โ the people who happen to be there are usually not typical.
Quota sampling โ the interviewer is told to fill fixed quotas (e.g. 20 men, 20 women), choosing subjects themselves. No sampling frame needed, but the interviewer’s choices introduce bias.
Cluster sampling โ split the population into naturally occurring clusters and randomly select whole clusters. Cheap for spread-out populations, but a cluster may not represent the whole.
Exam phrasing: when asked to “comment on the sampling method”, name one advantage and one disadvantage in context. “Surveying shoppers on a Tuesday morning is opportunity sampling โ quick and cheap, but people at work are excluded, so the sample is not representative of all shoppers.”
Quick check
Spot the bias
?A researcher surveys people leaving a gym about how much exercise the public gets. What is the flaw?
Calculate
Your turn โ a fair chance
5In a simple random sample of 20 from a population of 250, what is the probability that any one particular individual is selected? Give your answer as a decimal.
Hint: every member has an equal chance, so the probability is 20 ÷ 250.
Sampling · the data set
The OCR pre-release large data set
OCR issues a pre-release large data set (its name for the large data set) that you are expected to have worked with during the course. Questions in the exam may be set in its context.
You will not have a printout of the pre-release data set in the exam.
Any values you actually need for a calculation will be given to you in the question.
What is assessed is your familiarity with it: what its variables mean, what units they use, how the data was collected and sampled, and what its limitations are.
How to prepare properly: get the current data set from your teacher or the OCR website โ it is specific to your exam series. Explore it in a spreadsheet: find the outliers, notice which variables have missing entries, and ask yourself which sampling method produced it and what bias that might introduce. That is exactly the thinking the exam rewards.
Quick check
The data set in the exam
?What should you expect regarding the pre-release large data set in the exam?
Quick check
Define the sampling frame
?What is a sampling frame?
Sort it
Name that sampling method
Tap a description, then tap the sampling method it describes.
๐ฒ Simple random
๐ Systematic
๐งฑ Stratified
Match it
Statistical vocabulary
Tap a term on the left, then its definition.
Term
Definition
Recap
The big ideas to know
Population vs sample: a census covers everyone; a sample is a subset used to estimate
Sampling frame: the list you sample from โ random and systematic methods both need one
Simple random: equal chance for all; fair but needs a full frame
Systematic: k = N/n with a random start; fast, but beware periodic patterns
Stratified: proportional to stratum size โ check your parts add to the total
Non-random: opportunity and quota are quick but biased
Bias: a bigger biased sample is still biased โ size never fixes it
Data set: know OCR’s pre-release data set’s variables, collection and limitations โ not its values
That is OCR 2.01 โ the foundation every other statistics question stands on. Press Finish to see your score.
🏆
Mini-lesson complete!
⭐⭐⭐
You've worked through Statistical sampling for OCR A-level Mathematics A. 🎉
Your stars: 0 / 0
Next: test yourself in the Evaluate stage Confidence Quiz, then lock it in with Verify.
📣 Smashed it? Share your score
Challenge a mate to beat your stars, or show a parent how you got on.