memoza

The Memoza AcademyResearch

Mastery is more than a completion percentage

A green tick can hide very different evidence. What a platform actually observes, and why uncertainty belongs inside the estimate.

By Lucas Hsu · · 5 min read


Two students have a green tick beside the same topic. One answered two familiar questions correctly immediately after reading a solution. The other solved varied problems independently and returned successfully after a week.

The same icon could hide very different evidence.

A completion indicator can accurately report that an activity was finished. Interpreting it as durable knowledge requires additional assumptions about what was attempted, under which conditions, and how well that performance generalises.

We observe responses and infer capability

Knowledge is not directly visible to a learning platform. It observes actions: an answer, a sequence of steps, a request for help, or a decision to stop. The interpretation depends on the task.

A correct multiple-choice answer may reflect understanding, elimination of alternatives, or guessing. A slow derivation may reflect uncertainty, careful checking, accessibility needs, or an interruption. Timing can be useful, but it should not be treated as a transparent reading of competence.

Knowledge-tracing models address part of this inference problem by maintaining estimates of specific skills from sequences of responses (Corbett and Anderson, 1994). Their estimates are conditional on a model of learning and performance. They are not direct measurements of a student's mind.

For Memoza, that distinction should shape how an estimate is presented. A label such as “strong recent evidence” can be more defensible than an unexplained assertion that a concept is permanently mastered.

A small calculation shows the role of uncertainty

Consider an intentionally simplified model. Assume every attempt is an independent trial of the same fixed success probability. Start with a uniform Beta(1, 1) prior. Each correct answer adds one to the first parameter; each incorrect answer adds one to the second.

The resulting estimates differ from raw accuracy:

Observed resultsRaw accuracyPosteriorPosterior mean
2 correct from 2 attempts100%Beta(3, 1)75.0%
40 correct from 45 attempts88.9%Beta(41, 6)87.2%

Those figures are arithmetic consequences of this toy model. They are not Memoza scores, estimates of knowledge retained, or expected exam grades. They show why observed accuracy and an estimate that includes prior uncertainty need not coincide. They also show that two successful attempts do not justify certainty.

The assumptions are deliberately restrictive. Real students learn between attempts. Questions vary in difficulty. Repeated variants can share cues, making their outcomes dependent. Assistance changes what a correct answer demonstrates. Treating every response as an interchangeable trial can therefore produce misleading confidence.

More data help only insofar as the data and model support the inference being made.

Time changes the decision

A successful attempt today supplies evidence about performance today. Whether another check is worthwhile depends partly on when the capability will be needed again.

Settles and Meeder modelled recall using a learned memory half-life, connecting elapsed time and interaction history to predicted recall (Settles and Meeder, 2016). Their language-learning setting provides a useful example of making temporal assumptions explicit. It does not imply that a complex university skill has one objectively measurable half-life.

Even a well-calibrated estimate would need interpretation. A student preparing for tomorrow's exam may reasonably prioritise different reviews from someone trying to retain the course for next year's prerequisites.

The stopping question becomes: given the available evidence, is another activity on this objective worth the time right now? That is a decision under uncertainty, with a purpose and a horizon.

Transfer deserves separate evidence

Repeated success on a familiar format can leave generalisation untested. Can the student use the same idea with a different representation, combine it with another skill, or recognise a situation in which it should not be applied?

Butler's experiments found that repeated testing could improve transfer to new questions relative to repeated study (Butler, 2010). The finding supports the value of retrieval in those settings, while also highlighting that transfer has to be measured with tasks beyond simple repetition of the practised response.

For a dashboard, useful distinctions might therefore include coverage, recent independent performance, delayed performance, and evidence on unfamiliar applications. These describe different things. Combining them into a single number requires a justification that the interface should not hide.

Our preferred direction for Memoza is to connect every estimate to a useful next decision. A student should be able to understand why a review is recommended, why further practice can wait, and where evidence is still thin. A progress display earns its value by guiding those choices accurately.

How Memoza fits

Mastery in Memoza is not a stored score. It is folded out of an evidence log the database itself refuses to let the product update or delete, and the fold is pure enough that the whole record is recomputable by replaying that log from the beginning. What it computes is a posterior rather than a running accuracy, and old evidence decays back towards the prior as time passes, so a success from six weeks ago counts for less than one from yesterday without anybody editing it away. Beside the percentage, the same row prints how much evidence stands behind it: a phrase such as “measured once” and a separate grey chip reading “First look”, which is decided on the number of observations rather than on the smoothed value, so one answer cannot read as a verdict. Transfer is the part we have least of. The frozen evidence vocabulary has a value for a transfer success and the fold counts it as a field of its own, but the grading path every discipline uses today cannot emit that value, so it is a place kept for a measurement rather than a measurement.

References

Research types say how each source contributes to the argument. A predictive model does not carry the same evidential meaning as a randomised learning-outcome study.

  • Butler, A. C. (2010). Repeated testing produces superior transfer of learning relative to repeated studying. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(5), 1118–1133. https://doi.org/10.1037/a0019902 Experiments.
  • Corbett, A. T., & Anderson, J. R. (1994). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253–278. https://doi.org/10.1007/BF01099821 Model and empirical research.
  • Settles, B., & Meeder, B. (2016). A trainable spaced repetition model for language learning. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Volume 1, 1848–1858. https://doi.org/10.18653/v1/P16-1174 Computational research paper.

Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.

See what a few answers can and cannot show
  • Engineering · 5 min read

    Choosing the next question in Memoza

    Eligibility first, then the need behind the activity, then difficulty that serves the purpose. A recommendation you can reconstruct.

    Read guide →

Put the ideas to work.

Explore course practice in Memoza.

Try the demo, up to 3 questions

Explore course practice