Sampling
Stage 16 of 23 Strand 3 of 5 11 lessons
11 illustrated lessons, each teaching the why before the how.
Revise Sampling with flashcards →
Jump to a lesson
Populations and Samples #
A parameter is fixed; a statistic varies.
A parameter describes the whole population and a statistic describes only your sample
There are a thousand people in the population, and you can only ask thirty.
The population mean is fixed but unknown. You can compute the sample mean.
So a statistic is an estimate of a parameter, and it changes with each sample.
Now you
Is the true proportion of faulty parts a parameter or a statistic?
Is the proportion faulty in the box you opened a parameter or a statistic?
Lesson complete. Continue your journey in the app — your progress saves there.
Sampling Methods #
Chosen to avoid bias, not for convenience.
A sampling method is chosen to avoid bias rather than to be convenient
In a simple random sample, every member has the same probability of being chosen, and the selections are independent.
In a stratified sample, split the population into groups, then take from each group in proportion to its size.
In a systematic sample, take every tenth name on the list. It is fast, but it relies on the order of the list being fair.
A convenience sample takes whoever is easy to reach. At a gym, the venue decides who can be sampled.
Ask every member instead and it is a census — complete, and often impractical.
In a quota sample, fill a fixed count from each group, choosing freely. It is stratified sampling without the randomness.
Now you
A survey about exercise asks the first 20 people leaving a gym. What is wrong?
A school asks every single student about lunch. What is that called?
Lesson complete. Continue your journey in the app — your progress saves there.
Observational Studies and Experiments #
Assigning the treatment is what buys a cause.
Only a study that assigns the treatment at random can support a claim about cause
In an experiment the researcher assigns the treatment — here by a coin toss.
In an observational study you only record groups people put themselves into.
Self-chosen groups can differ before the treatment, and that alone can explain a gap.
Assigning at random spreads every other difference evenly over the two groups.
The control group is treated identically except for the treatment being tested.
Random assignment supports a cause; observing without it supports only an association.
Random sampling settles who a result covers; random assignment settles whether the result is a cause.
Now you
Volunteers were split by coin toss. What does that give the study?
An observational study finds a new drug goes with lower blood pressure. What may you conclude?
Lesson complete. Continue your journey in the app — your progress saves there.
Capture and Recapture #
The tagged share of a catch mirrors the pond.
Matching the tagged fraction of a catch to the pond estimates the whole population
Tag 20 fish and release them — they scatter through a pond of unknown size.
Recapture 30 fish: 6 carry tags — a fifth of the catch is tagged.
If the catch is representative of the pond, the fractions match — solving gives about 100 fish.
The estimate rests on assumptions: tagged fish mix fully, and every tag stays on.
Now you
Suppose the tagged fish spread right through the pond before the next catch. Is the estimate sound?
Suppose half the tags fall off before the next catch. Is the estimate sound?
Lesson complete. Continue your journey in the app — your progress saves there.
The Sampling Distribution #
Every possible sample mean, piled up.
Take many samples and their means form a distribution of their own
The population is ten scores with a mean of 5. One sample of four gives a mean of 6.
Take two more samples of four from the same ten scores, and each one has a different mean.
Keep only the means and they pile up between 4 and 6, where the scores ran 2 to 8.
Now you
Samples of 4 are taken from these ten scores again and again. How are the sample means spread?
Samples of 3 are taken from these ten scores again and again. How are the sample means spread?
Lesson complete. Continue your journey in the app — your progress saves there.
Mean and Variance of X-bar #
Unbiased, with variance divided by n.
The sample mean is right on average and its variance is divided by the sample size
On average the sample mean equals the population mean, so it is unbiased.
Independent variances add to , and dividing by n divides variance by .
So every extra observation narrows the estimate a little further.
Now you
The population variance is 4 and the sample size is 9. What is the variance of the sample mean?
The population variance is 25 and the sample size is 4. What is the variance of the sample mean?
Lesson complete. Continue your journey in the app — your progress saves there.
Standard Error #
Halve it and you need four times the data.
The standard error is the standard deviation of the sample mean, so it falls with the root of n
Take the square root of the variance and the n comes out as a root: .
The standard error falls fast at small n and then very slowly as n grows larger.
To halve the standard error you need four times as many readings.
Now you
Sigma is 12 and n is 4. What is the standard error?
Sigma is 6 and n is 25. What is the standard error?
Lesson complete. Continue your journey in the app — your progress saves there.
The Central Limit Theorem #
Means go normal whatever the population was.
Sample means become normal as the sample grows, whatever shape the population had
This population is lopsided and uneven, nothing like a bell.
Take means of samples from it and they pile up symmetrically anyway.
That is the central limit theorem, and it is why the normal curve appears so often.
Now you
As the sample size grows, what does the central limit theorem say about the population itself?
As the sample size grows, what does the central limit theorem say about the sample means?
Lesson complete. Continue your journey in the app — your progress saves there.
Applying the Theorem #
Thirty is a rule of thumb, not a sharp edge.
A sample of about thirty is usually enough for the normal approximation to be usable
Once the sample means are close enough to normal, every method for the normal distribution applies to them.
Thirty is a rule of thumb, not a sharp cutoff.
A badly lopsided population needs more than thirty, not exactly thirty.
Now you
Is a sample of 5 usually enough for the normal approximation?
Is a sample of 30 usually enough for the normal approximation?
Lesson complete. Continue your journey in the app — your progress saves there.
Approximating a Binomial #
Match the mean and variance to a normal.
A binomial with enough trials is close enough to a normal to be treated as one
Binomial bars already look bell-shaped once n is reasonably large.
Give the curve the same mean np and variance np(1 − p) as the bars, and it sits on top of them.
The fit needs room in both tails: np and n(1 − p) must both be greater than 5.
Now you
X ~ . Which normal approximates it?
X ~ . Which normal approximates it?
Lesson complete. Continue your journey in the app — your progress saves there.
Continuity Correction #
Bars have width; a curve does not.
Swapping bars for a curve needs half a unit of width added back at each edge
A binomial bar for X = 7 really covers everything from 6.5 to 7.5.
The curve has no bars, so you must say where the bar edges were.
keeps the whole bar for 7, so the boundary is its left edge, 6.5.
P(X > 7) drops that bar, so the boundary moves out to its right edge, 7.5.
Now you
Approximating P(X < 7), which boundary do you use?
Approximating , which boundary do you use?
Lesson complete. Continue your journey in the app — your progress saves there.