A sample of observations has sample mean and sample standard deviation .
Construct a confidence interval for the population mean .
Choose the right interval formula. The population standard deviation is unknown and has been estimated from the sample, so the correct construction is the -interval
with degrees of freedom. With the sample is large enough that the central limit theorem makes the interval valid even if the underlying data are not normal — which matters here, because is nearly five times the mean, a strong hint of a heavily skewed variable.
Compute the standard error of the mean. This measures how much itself would bounce around from sample to sample:
The in the denominator is the reason a large sample helps: the raw spread of is cut by a factor of nearly .
Look up the critical value for confidence. For confidence, and each tail carries . With ,
This is only a whisker above the normal value , because a distribution with hundreds of degrees of freedom is almost indistinguishable from the standard normal. Using here is a defensible shortcut, but is the technically correct choice whenever is estimated.
Multiply to get the margin of error.
Round only at the very end; rounding to first would shift the endpoints in the third decimal place.
Assemble the interval.
Reported to two decimals, the confidence interval for is . For comparison, the -based interval would be — the same to within two hundredths, as expected at this sample size.
Interpret it correctly. The statement is about the procedure, not about this one interval: if the sampling were repeated many times, of the intervals built this way would contain the true mean . It is not correct to say that of the individual data values fall in — the data spread is governed by , roughly seventeen times wider than this interval, which describes only the uncertainty in the mean.
Need to solve a different problem like this? Open the solver →