At a glance
- Course
- Maths AI, current guide (first assessment 2021)
- Calculator
- GDC required in every AI paper
- Core tests
- Pearson's r, Spearman's rank, chi-squared
- Default accuracy
- 3 significant figures unless told otherwise
- Course runs to
- November 2028; a new course is first assessed in 2029
Which statistical tool answers which question
| Question in the paper | Tool | What you write down |
|---|---|---|
| How strong is the linear relationship? | Pearson's r | Value of r, then strength and direction in words |
| Predict y from x | Regression line of y on x | y = ax + b with values to 3 s.f., then the prediction |
| Is there a monotonic relationship, or is the data ranked? | Spearman's rank, r_s | Value of r_s and what it says about the ranks |
| Are two categorical variables independent? | Chi-squared test for independence | Hypotheses, degrees of freedom, statistic or p-value, conclusion |
| Probability for a continuous, symmetric variable | Normal distribution | Distribution with mean and standard deviation, then the probability |
HL adds further statistics, including more hypothesis testing and confidence intervals. Check your school's topic list for the exact content at your level.
Worked example 1: regression line and Pearson's r
Data: x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5. Find the regression line of y on x and r.
- Enter x and y as two lists on the GDC and run linear regression. It returns y = 0.6x + 2.2 and r = 0.775 (3 s.f.).
- Check by hand: the means are 3 and 4. The sum of (x - 3)(y - 4) is 4 + 0 + 0 + 0 + 2 = 6, and the sum of (x - 3)^2 is 10, so the gradient is 6/10 = 0.6 and the intercept is 4 - 0.6 × 3 = 2.2.
- r = 6/sqrt(10 × 6) = 6/sqrt(60) = 0.7746, which rounds to 0.775. The sum of (y - 4)^2 is 4 + 0 + 1 + 0 + 1 = 6.
- Interpret: a fairly strong positive linear correlation. Predict y when x = 3.5: 0.6 × 3.5 + 2.2 = 4.3. This is interpolation, inside the data range, so it is reasonable.
- Do not predict for x = 20. That is extrapolation, and examiners expect you to say the prediction is unreliable.
Worked example 2: Spearman's rank
Five students score 56, 72, 64, 88, 80 in maths and 60, 70, 75, 85, 78 in physics. Find r_s.
- Rank each list from lowest to highest. Maths ranks: 1, 3, 2, 5, 4. Physics ranks: 1, 2, 3, 5, 4.
- On the GDC, enter the ranks as two lists and find r for the ranks. That value is r_s = 0.9.
- Check with the formula: the rank differences d are 0, 1, -1, 0, 0, so the sum of d^2 is 2. r_s = 1 - 6 × 2/(5 × (25 - 1)) = 1 - 12/120 = 0.9.
- Interpret: a strong positive agreement between the rankings. Spearman's rank detects any consistently increasing relationship, not only a straight line, which is why it suits ranked or curved data.
Worked example 3: chi-squared test for independence
100 students are asked whether they prefer studying in the morning or the evening. In group X, 30 prefer morning and 20 prefer evening. In group Y, 20 prefer morning and 30 prefer evening. Test at the 5% level whether preference is independent of group.
- Hypotheses. H0: study preference is independent of group. H1: study preference is not independent of group.
- Expected frequencies: each row total is 50 and each column total is 50, so each expected value is 50 × 50/100 = 25. All expected values are at least 5, so the test is valid.
- Degrees of freedom: (rows - 1)(columns - 1) = 1 × 1 = 1.
- Statistic: each cell gives (O - E)^2/E = 25/25 = 1, so chi-squared = 4. Entering the observed table in the GDC gives chi-squared = 4 and p = 0.0455.
- Conclusion: p = 0.0455 is less than 0.05, so reject H0. There is evidence at the 5% level that study preference depends on group. Using the critical value instead: 4 > 3.841 gives the same decision.
Worked example 4: the normal distribution
Test marks are modelled by X ~ N(65, 10^2).
- P(X > 80): use the normal cumulative function with lower bound 80, a very large upper bound, mean 65 and standard deviation 10. Answer 0.0668 (3 s.f.). Check: z = (80 - 65)/10 = 1.5, and P(Z > 1.5) = 0.0668.
- The mark needed to be in the top 10%: use the inverse normal with area 0.9 to the left. The GDC gives 77.8. Check: z = 1.2816, and 65 + 1.2816 × 10 = 77.8.
- Always write the distribution and the probability statement, for example P(X > 80) = 0.0668. A bare number with no working risks losing the method mark if the answer is wrong.
Common mistakes
- Confusing standard deviation and variance. N(65, 10^2) means the standard deviation is 10; entering 100 into the calculator gives a wrong answer.
- Writing 'correlation proves causation'. A strong r shows association only.
- Stating H0 as 'the variables are dependent'. H0 is always independence.
- Comparing the chi-squared statistic with the p-value, or the p-value with the critical value. Compare like with like.
- Rounding early. Keep full calculator values until the final answer, then give 3 significant figures.
- Using the regression line of y on x to predict x from y.
Exam technique, and how one-to-one lessons help
Most AI statistics questions reward the same shape of answer: name the test or model, show the calculator inputs (lists, hypotheses, degrees of freedom), give the output to the right accuracy, then interpret it in one sentence about the real context. The interpretation mark is the one students most often drop, especially in Paper 2.
A tutor who teaches AI works with the student's own calculator model on screen, so every routine is practised exactly as it will be used in the exam. Lessons pair a short concept explanation on the shared whiteboard with past-paper questions where the student talks through each choice, and the tutor keeps a list of recurring slips, such as variance entered as standard deviation, until they stop happening.
Self-check
- For x = 2, 4, 6, 8 and y = 3, 7, 9, 13, find the regression line of y on x. (Answer: y = 1.6x, the intercept is 0)
- Ranks 1, 2, 3, 4 and 2, 1, 4, 3: find r_s. (Answer: 0.6)
- A 3 by 2 contingency table: how many degrees of freedom? (Answer: 2)
- X ~ N(50, 4^2): find P(X < 46). (Answer: 0.159)
- Explain in one sentence why predicting far outside the data range is unreliable.
Common questions
Do I need to calculate r or chi-squared by hand in IB Maths AI?
The GDC is expected for the calculations, but you must know what to enter, which hypotheses you are testing, how degrees of freedom are found and how to interpret the output. Understanding the hand method helps you spot a wrong entry.
When should I use Spearman's rank instead of Pearson's r?
Use Spearman's when the data is ranked, when the relationship is consistently increasing or decreasing but not linear, or when outliers distort Pearson's r. Pearson's r measures linear correlation only.
What happens if an expected frequency is below 5?
The chi-squared test is not reliable. In exam questions you may be asked to combine rows or columns so that every expected frequency is at least 5, which also changes the degrees of freedom.
How accurate should my answers be?
Unless the question says otherwise, give exact answers or answers correct to three significant figures. Keep full calculator values in intermediate steps.
How much do IB Maths AI lessons cost?
$15 a lesson, one flat rate for every subject and level. Families choose a weekly plan of 1 to 5 lessons billed monthly. Lessons are 60 minutes, one to one and online, and the first lesson is a free trial.
Sources
Dates and figures on this page come from these official and published sources. Always confirm deadlines on the official page before acting on them.