Research note 001: when to go deeper, when to move on
A proposed experiment on practice sequencing: deepen one concept or move across the syllabus? What we built to test it, and what we do not know yet.
By Lucas Hsu · · 4 min read
There is a decision hiding inside every practice session, and it is made again after every question: stay on this concept and go deeper, or move to something else. Tutors make it by feel. Students make it by default, usually by finishing whatever chapter they are in. Memoza makes it with a scoring function. None of us can currently prove our answer is right.
This is a research note, not a results report. It describes an experiment we have designed and built the machinery for, and have not run. Nothing in it is evidence that Memoza improves outcomes.
The question
The companion guide argues that strong exam preparation is breadth, then depth, then breadth again. That argument is a synthesis: the laboratory literature on one side, the logic of a marked exam paper on the other. The laboratory result it leans on is real — in controlled studies, interleaving related problem types tends to beat blocking them, even though blocked practice feels more productive while it happens. But laboratory tasks last minutes and a course lasts months, and a result that survives the distance between those two is rarer than either.
So the honest position is this: we believe ordering matters, we believe the right order is roughly breadth, depth, breadth — and we have not measured any of it in our own product.
What we built to test it
Memoza's recommender scores every concept in a course and serves the one that needs practice most. To test sequencing, we did not build a second recommender. We built a stage that can restrict which candidates the existing one may choose from, without changing how it scores them. Four policies, and one of them is deliberately nothing:
- Balanced — the identity. No restriction, exactly the recommender every student has today.
- Depth first — stays in the current concept's neighbourhood until the evidence clears a mastery bar or the run reaches its limit.
- Breadth first — moves across topics before returning to deepen any of them.
- Adaptive — deepens where the evidence says depth is owed, and broadens otherwise.
The property that makes a comparison meaningful: a policy restricts the order, never the menu. Every arm chooses from the same candidate pool, so a difference between arms is a difference in sequencing and nothing else. And when a policy cannot do what it wants — there is no admissible candidate to deepen into — it falls back and records that it fell back, so the log distinguishes "chose to move on" from "could not go deeper".
The experiment we propose
Assignment would be random, derived by hashing so that an auditor can re-derive every allocation and verify nothing was hand-placed. Participation is by explicit consent, and only consenting students are ever assigned to an arm.
A student outside the study is not a control group.
That sentence is enforced in the export, not just promised: students who never consented appear with an empty cell where an arm would be, never as "control". Withdrawing consent deletes the assignment, and erasure works on every table this study touches.
Outcomes would be measured by assessments that sit outside the practice loop — the same paper before, after, and later, never feeding the mastery model it is trying to measure — and, where a student separately consents, by the mark their university actually gave them.
Everything above is a proposal. The numbers inside the policies — how long a depth run may go, what mastery bar ends one — are starting points that a preregistration freezes before the first participant. They are not knobs to turn while the study runs.
What we cannot measure yet
Four gaps, recorded before the study rather than discovered in the analysis:
- "Related concept" currently means "same topic". The course graphs we have author prerequisite relationships almost exclusively, so a depth policy's neighbourhood is coarser than the word "related" suggests.
- Transfer is not tested. Whether depth on one concept helps a neighbouring one is exactly the interesting question, and we have no question type built to ask it yet. We reserved the slot rather than faking the measurement.
- Depth runs get interrupted structurally. The rotation between activity types outranks the sequencing policy, and in one course's template mix roughly one step in eight cannot continue a depth run. Those steps are stamped as fallbacks, so the analysis can separate "the policy chose to broaden" from "the policy had no move".
- Long-term stability is not a covariate. We declared it unavailable rather than substituting a different quantity that happens to be lying around.
Where this stands
Stage: proposed experiment. The machinery is live in production and deliberately doing nothing — every student today gets the balanced policy, and the other three are unreachable until a study assigns them. Before a first participant exists, four things must: an approved preregistration, a decision-register entry, a privacy impact addendum, and a participant information sheet. If we run the study, the design is frozen and published before the first participant, and the result gets published whether or not it flatters us.
Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.
See how Memoza sequences your practice
