Half of the AP Statistics exam is six free-response questions, and they are where scores separate. The multiple-choice section rewards recognition; the free-response section rewards precise writing, and most students have never been taught what "precise" means to an AP Statistics reader. This guide covers the structure, the scoring mechanics, the templates that reliably earn points, and a complete worked example marked against a rubric.
(AI-Math is independent and not affiliated with the College Board. The problems below are original; released questions and official scoring guidelines are published on the College Board's website and are free.)
The structure of the free-response section
Ninety minutes, six questions, 50% of the exam.
- Questions 1тАУ5 are shorter, roughly 12тАУ13 minutes each, and together account for about three quarters of the section's score. They typically cover: exploring data, study design, probability or random variables, and inference тАФ usually two inference questions among them.
- Question 6 is the Investigative Task, roughly 25 minutes, worth about a quarter of the section on its own. It introduces something you have not been taught тАФ a modified statistic, an unfamiliar procedure, a simulation-based approach тАФ and asks you to reason with it using what you do know. It is not a trick; it is a test of whether you understand statistics rather than having memorised procedures.
How scoring actually works
Each question is scored 0 to 4, holistically:
| Score | Meaning |
|---|---|
| 4 | Complete response |
| 3 | Substantial response |
| 2 | Developing response |
| 1 | Minimal response |
| 0 | No credit-worthy work |
Within a question, each part is first labelled essentially correct (E), partially correct (P) or incorrect (I). Those labels combine into the 0тАУ4 score тАФ for a three-part question, roughly, three Es gives a 4, two Es and a P gives a 3, and so on. The exact combination rules are published with each question's scoring guidelines.
Three consequences worth internalising:
- Communication is graded, not just computation. A correct number with no interpretation is regularly a P.
- Errors do not always cascade. If part (a) is wrong but you use your wrong value correctly in part (b), part (b) can still be essentially correct. Never abandon a question mid-way.
- Extra wrong statements can cost you. Readers mark what you wrote. If you give a correct conclusion and then add an incorrect one, you can be marked down for the contradiction. Write one clear answer, not three hedged ones.
The templates that earn points
Describing a distribution
Always four elements: shape, centre, spread, unusual features тАФ in context.
"The distribution of commute times is skewed right, centred near a median of 24 minutes, with an interquartile range of about 11 minutes, and there is one apparent outlier above 70 minutes."
Confidence interval
- Name the procedure and the parameter in context.
- Check conditions with evidence.
- Calculate the interval, showing the values used.
- Interpret: "We are 95% confident that the true [parameter in context] is between [a] and [b]."
Never write "there is a 95% probability the parameter is in this interval." The parameter is fixed; the interval is random. That sentence is the most common single error on the exam.
Significance test
- Hypotheses about a parameter, defined in words: ", , where is the proportion of all customers at this store who use the app."
- Procedure name and conditions verified.
- Test statistic and p-value.
- Conclusion: compare the p-value to , state reject or fail to reject, and give the conclusion in context.
"Because the p-value of 0.031 is less than , we reject . We have convincing evidence that more than 40% of customers at this store use the app."
Note "fail to reject", never "accept". Our p-value calculator and hypothesis test calculator handle step 3; steps 1, 2 and 4 are the graded writing.
Interpreting a slope
"For each additional hour studied, the predicted exam score increases by 3.4 points." Predicted, in context, with units тАФ all three are required.
A full worked FRQ
A school district wants to know whether a new 8-week tutoring programme improves scores on a standardised maths test. Sixty volunteers are randomly assigned, 30 to the programme and 30 to a control group that receives no tutoring. After 8 weeks, the tutored group has a mean score of 78.4 with standard deviation 9.2; the control group has a mean of 73.1 with standard deviation 10.1. Both distributions are roughly symmetric with no outliers.
(a) Is this an experiment or an observational study? Justify.
It is an experiment, because the researchers imposed a treatment тАФ assignment to tutoring or no tutoring тАФ rather than merely observing existing groups, and the assignment was made at random.
(b) Carry out an appropriate test at .
Hypotheses. against , where is the mean score for all students like these if tutored and is the mean score if not tutored.
Procedure and conditions. A two-sample t-test for a difference of means. Random: subjects were randomly assigned to the two groups. Independence: random assignment makes the two groups independent, and each group's responses are independent of the others'. Normality: both sample sizes are 30 and both distributions are stated to be roughly symmetric with no outliers.
Computation. The standard error of the difference is
With roughly 57 degrees of freedom, the one-sided p-value is about 0.019.
Conclusion. Because , we reject . We have convincing evidence that the tutoring programme increases mean standardised test scores for students like those in this study.
(c) Can the district conclude that the programme would raise scores for all students in the district? Explain.
No. The subjects were volunteers, not a random sample of district students, so the causal conclusion is valid only for students like these volunteers. Random assignment supports a causal claim within the study; random selection is what supports generalisation to a wider population, and it is absent here.
How this would be marked. Part (a) is E if it both names "experiment" and justifies with imposed treatment. Part (b) is scored across four components тАФ hypotheses, conditions, mechanics, conclusion in context тАФ and dropping the context in the conclusion or listing conditions without evidence drops it to P. Part (c) is E only if it distinguishes random assignment from random selection; answering "no, because the sample is too small" would be incorrect.
Part (c) is the discriminating question. It is testing one idea тАФ assignment supports causation, selection supports generalisation тАФ and that idea appears on the exam almost every year in some form.
The most expensive habits
- Naming conditions without checking them. "Large counts" is not a condition check; " and " is.
- Hypotheses about statistics. is wrong, always. Hypotheses concern parameters.
- Conclusions with no context. "Reject the null" alone caps the part at partially correct.
- Accepting the null. Failing to find evidence is not evidence of no difference.
- Answering more than was asked. A contradictory extra sentence can cost the point you already earned.
- Blank parts. A partially reasoned answer often earns something; a blank never does. On the Investigative Task especially, write down the reasoning even if you cannot finish.
How to practise
Many years of released free-response questions with official scoring guidelines and samples of scored student work are published free. Use them like this: do one question under time, then mark your own answer strictly against the guideline, component by component, awarding only what the rubric literally requires. Then rewrite the answer to a 4.
Doing that ten times is worth more than a hundred multiple-choice questions, because the gap between what you would give yourself and what a reader would give you is your score gap.
Related: AP Statistics exam guide and Is AP Statistics hard?