memoza

The Memoza AcademyResearch

Can adaptive breadth and depth improve learning

Recognising a learner's state is not the same as responding well. What a fair comparison with fixed schedules would have to look like.

By Lucas Hsu · · 5 min read


A student struggling with a prerequisite may benefit from focused work. A student who can execute several methods but confuses their applications may benefit from mixed practice. A student returning after a delay may need to review something that looked secure last week.

It is tempting to conclude that a system which recognises these states must outperform a fixed schedule. That conclusion is premature. The system has to identify the states accurately, choose an effective response, and justify the time spent collecting the information.

Memoza's research hypothesis is that a policy which adjusts breadth, depth, and review to learner evidence can improve outcomes relative to credible fixed policies. The word “can” matters: this is a question to test, not a result we have demonstrated.

Define the competitors precisely

The informal labels BFS and DFS are insufficient for an experiment. We need executable schedules. BFS and DFS as a way to think about learning sets out what the analogy does and does not claim.

A candidate comparison could define the policies as follows:

PolicyAllocation rule
Fixed depth-orientedPredetermined blocks within each eligible topic before moving to the next
Fixed breadth-orientedA predetermined rotation across eligible topics
AdaptiveAllocation changes with learner evidence, within the same content and eligibility constraints

All three would need explicit rules for prerequisites, review opportunities, feedback, and stopping. If the depth-oriented arm alone requires students to pass a mastery threshold before moving on, it is already adapting to performance. That may be a worthwhile comparator, but it should not be described as a purely fixed schedule.

The controls should also be credible study strategies. Preventing the fixed arms from receiving any spaced review would make it difficult to learn whether the adaptive decision itself adds value.

What the existing evidence justifies

There is reason to investigate state-dependent allocation. The interleaving literature shows variation across learning materials and conditions, rather than a uniform benefit from mixing everything (Brunmair and Richter, 2019).

Personalised review has also outperformed common review schedules in a semester-long language-learning study (Lindsey et al., 2014). That finding makes adaptation plausible, while leaving the broader university sequencing question open.

Tabibian and colleagues developed and evaluated an approach to optimising spaced review using a model of memory (Tabibian et al., 2019). Their work supplies a concrete precedent for treating scheduling as an optimisation problem. Its objective and assumptions are more specific than the full choice among prerequisite repair, comparison, transfer, and review.

These strands support a research programme. Combining them in a product creates a new intervention, whose effects cannot be obtained by adding together results from the underlying papers.

Adaptation has costs and failure modes

An adaptive policy can overreact to a careless error. It can mistake a difficult item for a weak learner. It can spend too much time diagnosing, or repeatedly select a familiar question because its own model predicts success there.

It can also neglect curriculum coverage. A local decision to repair one more weakness may look reasonable each time, yet leave several topics untouched by the exam. Useful sequencing needs some account of the remaining time and the course as a whole.

The graph itself may be wrong. An overly strict prerequisite relationship can block a student from material they are capable of learning. An overly loose relationship can send them into a task whose difficulty comes from missing foundations.

These are reasons to compare a first adaptive policy with a strong simple baseline. Simplicity has a practical advantage when additional complexity produces no meaningful improvement.

Decide in advance what would change our view

For a fixed amount of study time, the most direct claim would concern performance on an independent assessment. A convincing benefit should survive a delayed test if the product claims to improve retention.

Several results would weaken the proposed policy. Either fixed strategy might outperform it. The adaptive group might complete more questions without scoring better. It might improve immediate performance while losing the advantage after a delay. It might help one course and impair another.

An imprecise null result requires care. A small study that cannot distinguish a useful gain from a useful loss has not established equivalence. Conversely, a precise result that excludes the benefit needed to justify the system's complexity would be a strong reason to revise the design.

Failure of one policy would not disprove every possible adaptive policy. It would show that this policy, under these conditions, did not earn its advantage. Similarly, a positive average result would not prove that each individual switching decision was optimal.

A small study that cannot distinguish a useful gain from a useful loss has not established equivalence.

That is the level at which the hypothesis becomes useful. We can specify a policy, compare it fairly, and change it when the evidence warrants. The breadth and depth metaphor has done its job only when it leads to a decision we are willing to test.

How Memoza fits

The seam this comparison would run in is built, and it is deliberately inert. Four policies are named in the recommender, balanced, depth first, breadth first and adaptive, and a policy restricts which of the eligible candidates the scorer is allowed to choose between rather than replacing the scorer. Balanced is the identity: it restricts nothing, and it is what a session resolves to unless a running study has enrolled that student in an arm, which none has. The machinery a study would need is there too, as studies, arms, consents and assignments, with an assignment keyed to the course family rather than to a particular version so that a version bump does not re-randomise a participant halfway through. What runs today is the control arm, and the record each step leaves, which policy governed it and how many candidates it was choosing between, is what a later comparison would be read from.

References

Brunmair, M., & Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin, 145(11), 1029–1052. https://doi.org/10.1037/bul0000209 Meta-analysis.

Lindsey, R. V., Shroyer, J. D., Pashler, H., & Mozer, M. C. (2014). Improving students' long-term knowledge retention through personalized review. Psychological Science, 25(3), 639–647. Classroom study.

Tabibian, B., Upadhyay, U., De, A., Zarezade, A., Schölkopf, B., & Gomez-Rodriguez, M. (2019). Enhancing human learning via spaced repetition optimization. Proceedings of the National Academy of Sciences, 116(10), 3988–3993. Computational and empirical research.

Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.

See the sequencing that runs today

Put the ideas to work.

Explore course practice in Memoza.

Try the demo, up to 3 questions

Explore course practice