Research9 min read

How AI tracks your child's mastery (and why it is better than marks)

Bayesian Knowledge Tracing, spaced repetition, and SmartScore explained in plain language — what the numbers mean, why percentage scores mislead parents, and what to look for in a mastery report.

The problem with percentage marks

Most parents grew up with marks. 72% in the Friday maths test. 8/10 on the spelling quiz. These numbers feel informative because they are familiar, but they have a fundamental flaw: they tell you what happened on one day, not what your child actually knows.

A child who scores 80% on a multiplication test on Friday may have forgotten 40% of it by the following Wednesday. A child who scores 60% may have genuinely understood the core concept and made careless errors on the rest. A mark cannot tell you which situation you are in, because a mark measures performance, not learning.

This distinction — between performance and learning — is not semantic. It has been replicated across hundreds of studies in cognitive science since the 1980s. The field calls it the "performance-learning distinction," and it is one of the foundational reasons why adaptive learning systems have moved away from percentage scores.

What Bayesian Knowledge Tracing actually is

Bayesian Knowledge Tracing (BKT) was introduced by Corbett and Anderson in 1995 and has been refined extensively since. The core idea is straightforward: instead of asking "what score did this child get today?", BKT asks "given everything I have observed about this child's responses over time, what is the probability that they have genuinely acquired this knowledge?"

The model tracks four hidden parameters for each skill:

  • P(L₀) — the probability the child already knew this skill before any instruction
  • P(T) — the probability of learning the skill after a single learning opportunity
  • P(S) — the probability of making a slip (getting a question wrong despite knowing the skill)
  • P(G) — the probability of guessing correctly (getting a question right despite not knowing the skill)

After each response — correct or incorrect — the model updates its estimate of whether the child truly knows the skill. This updated probability is called pMaster (probability of mastery).

Crucially, pMaster accounts for guessing and slipping. If your child answers three multiple-choice questions correctly, a raw score says 100%. But if each question had a 25% guess probability, BKT recognises that some of those correct answers may be lucky guesses and will not inflate pMaster to 1.0. Conversely, if your child who clearly understands a concept makes one careless error, BKT will not dramatically drop their mastery estimate the way a score would.

The difference between pMaster and confidence

In more sophisticated implementations, mastery tracking reports two numbers:

  • pMaster — the current estimated probability that the child has acquired the underlying skill
  • Confidence — how much evidence the model has collected to make that estimate

A child who has answered two questions correctly will have a positive pMaster, but low confidence — there is not enough data to be sure. A child who has answered twenty questions with consistent accuracy will have both high pMaster and high confidence. The combination matters: a high pMaster with low confidence tells you to ask a few more questions before concluding mastery; a high pMaster with high confidence tells you this skill is genuinely acquired and it is time to move on.

When reading a mastery report, look for both values. A single pMaster number without a confidence indicator is incomplete.

How spaced repetition scheduling works

Knowing that a child has mastered a skill today does not mean they will retain it. Memory research, going back to Ebbinghaus in 1885 and expanded dramatically in recent decades, shows that retention follows a predictable decay curve. Without review, even well-learned material fades.

Spaced repetition scheduling (SRS) uses this decay curve to schedule reviews at the optimal moment — just before a child is about to forget. The most influential modern SRS algorithm is FSRS (Free Spaced Repetition Scheduler), version 2023, which is trained on real learner data and adapts to individual forgetting rates.

In practice, this means:

  • A skill a child learned yesterday might be reviewed tomorrow
  • A skill they reviewed twice with strong recall might not appear again for three weeks
  • A skill they consistently struggle with will appear more frequently until the system sees consistent correct responses

From a parent's perspective, this looks like the AI tutor occasionally returning to something your child "already knows." This is intentional. The revisit is timed to reinforce the memory trace at the moment it is strongest, which is counterintuitively not immediately after learning. Research shows that reviewing material at the point of near-forgetting, rather than when it is fully fresh, produces significantly stronger long-term retention.

What to look for in a SmartScore report

A SmartScore is a composite metric that combines pMaster, recency of last successful recall, and the number of successful spaced-repetition reviews completed. It gives you a single number (0–100) that represents how confidently a child has durable command of a skill, not just whether they got it right recently.

When reading a SmartScore report for your child:

Look for clusters of low scores in the same strand. If every Measurement descriptor in Year 3 Maths is below 50, that is a signal that this strand needs deliberate attention, not just more general practice.

Distinguish low scores with high confidence from low scores with low confidence. A skill with pMaster 0.4 and 15 data points is genuinely struggling. A skill with pMaster 0.4 and 2 data points just hasn't been visited enough — schedule a lesson on it.

Check the "last reviewed" date. A skill with a high pMaster but a last review date of six months ago may have decayed. The spaced repetition system should have flagged this for review; if it hasn't appeared, it may be worth manually scheduling.

Use the report for your compliance document. A printout showing curriculum codes, pMaster scores, confidence levels, and review history is compelling evidence for a VRQA, NESA, or IHIP submission. It demonstrates not just that your child was exposed to the material, but that the system can verify they retained it over time.

Why this is better than marks — the summary

Percentage marks measure a single performance event. BKT-based mastery tracking measures the accumulation of evidence about durable learning. Spaced repetition scheduling ensures that assessed knowledge is tested at the moment of near-forgetting, producing retention data rather than performance data. The combination gives you a picture of what your child genuinely knows — not just what they could do last Friday morning when the test happened to occur.

For home educators, this is particularly valuable because you are not constrained by a fixed Friday test schedule. You can use the mastery data to decide when to move on, when to revisit, and when to enrich — decisions that in a classroom have to be made for thirty children at once but in a home setting can be made for one.

Related posts

Research

FSRS-23 explained: how Docent schedules your child's reviews

A plain-English guide to FSRS-23 — the spaced-repetition algorithm behind Anki 23+ and Docent — covering the Forgetting Curve, why cramming fails, how FSRS differs from BKT mastery tracking, and what optimal review scheduling looks like for a Year 4 Maths student.

Read more →
Built by Innovenses Pty Ltd · Melbourne. Docent is an AI tool; it supports — and never replaces — a parent or qualified teacher. Privacy · Terms · Blog · System status.