How to test whether Memoza improves learning
Engagement is not evidence. The study design that could answer whether our sequencing helps, and what would change our mind.
By Lucas Hsu · · 6 min read
A student answers more questions after joining Memoza. They come back on more days, their mastery estimates rise, and they report feeling more prepared.
These observations matter for usability and engagement. They leave the central educational question open: would the student have learned more, less, or the same amount with another way of spending that time?
This article proposes an evaluation of Memoza's sequencing policy. It describes a study we could run, rather than a completed trial, a registered protocol, or an announced university partnership.
Start with a claim the study can identify
The initial question should be narrow: does the proposed adaptive policy improve learning compared with fixed breadth-oriented and depth-oriented schedules when the study conditions are otherwise comparable? What that comparison is worth, and what it cannot settle on its own, is the subject of can adaptive breadth and depth improve learning; this piece is about how one would actually run it.
That comparison estimates the added value of sequencing inside the product. It would not establish the value of the entire Memoza experience compared with ordinary studying, another platform, or private tutoring. Those require additional comparisons.
One candidate design would use a defined course unit, a common study-time allowance, and an independent delayed assessment. The primary outcome could be the score on that assessment seven days after the final assigned session. The interval and study duration would need to be finalised before recruitment, with the course calendar taken into account.
Randomise the policy and preserve the assignment
Students could be allocated to one of three conditions: a predetermined depth-oriented schedule, a predetermined breadth-oriented schedule, or the adaptive policy. Randomisation could be stratified by course and baseline performance to improve balance on variables chosen in advance.
Each student's assignment would persist throughout the intervention. Switching people between arms in response to their progress would compromise the original comparison unless that switching were itself part of a prespecified design.
If students share practice closely within tutorial groups, assigning whole groups may be more appropriate. That changes the sample-size calculation and analysis because students in the same group cannot be treated as independent observations.
Change sequencing while keeping the teaching conditions comparable
All arms should draw from the same eligible question bank and use the same explanations, feedback, interface, and grading rules. They should receive comparable access and incentives. The policy versions should be fixed for the main experiment.
The exact sequence of questions will differ; that is part of the intervention. The content available to each arm should not differ simply because one group receives better-authored material.
Study-time limits also need an operational definition. Active interaction time, elapsed time, and time spent outside the platform are different quantities. A controlled study can standardise assigned session lengths, while records and brief reports can document departures from the plan.
Test capabilities on independent material
The proposed measurement sequence is a baseline assessment, the assigned study sessions, an unseen immediate post-test, and an unseen delayed test. Actual course-exam results could be a secondary outcome where students consent and access is available.
“Unseen” must mean more than new numbers. Held-out tasks should test the same objectives without being trivial variants whose answers can be reproduced from a memorised template. Appropriate transfer tasks can ask for a new representation, explanation, or application while remaining within the course's expected scope.
Butler's work on retrieval and transfer illustrates why assessment beyond the practised response is informative (Butler, 2010). A study should specify how far its transfer tasks differ from practice, rather than use “transfer” as an unrestricted claim.
The assessment and rubric should be reviewed independently of the sequencing team where feasible. Markers should be unaware of students' assigned condition. The product's own mastery estimate should not be the primary endpoint: a policy might raise that estimate without improving external performance.
Prespecify the analysis and the useful effect
Before recruitment, the protocol should define the primary outcome, the two main adaptive-versus-control comparisons, the analysis model, exclusions, and treatment of missing outcomes. A baseline-adjusted analysis could estimate the differences between groups, with course or group structure included as appropriate and multiplicity addressed for the planned comparisons.
The main analysis should preserve randomised assignment. Comparing only students who completed the whole schedule risks selecting different kinds of learners across arms. Missing assessments need a stated handling strategy and sensitivity analyses; they should not silently disappear or automatically become zero scores.
Lakens' guidance on sample-size justification emphasises specifying what effect would be informative and what precision a study needs (Lakens, 2022). For Memoza, the threshold should reflect a meaningful educational gain or a worthwhile saving of study time. A convenient number of beta users is a recruitment fact, not a power calculation.
A convenient number of beta users is a recruitment fact, not a power calculation.
Efficiency claims need their own design. Dividing score gain by self-selected study hours can be misleading, because the intervention may change how long students choose to work. One study can compare scores under a common time budget. Another could compare time to an independently defined performance criterion, with retention checked afterwards.
Preregistration makes the distinction between planned tests and later exploration visible (Nosek et al., 2018). It does not guarantee a good design, but it helps readers judge whether the reported analysis was chosen after seeing the results.
Make the result credible enough to disagree with
Participation should be voluntary, research use of data should be explained, and institutional ethics review should be obtained where required. The authors' ownership of Memoza should be disclosed. External academic involvement should be described accurately if it is secured.
A credible protocol should also specify how unfavourable and inconclusive findings will be reported. Code, task specifications, and sufficiently protected data can support scrutiny where sharing is appropriate.
The first study need not prove that Memoza works everywhere. It should provide an honest answer about a defined policy, learner population, and course setting. That answer would give us something firmer than engagement metrics on which to base the next decision.
How Memoza fits
Some of that protocol is already substrate rather than intention. The recommender holds the four named policies such a comparison would need, and balanced, the one that restricts nothing, is a control arm by name rather than the absence of an assignment. A study cannot be recorded without a participant notice version and a path to its preregistration; a granted consent row is read before any assignment is written; and an assignment is keyed on the course family, so a version bump cannot re-randomise somebody in the middle of a study. Nobody has been assigned to anything, because no study has been opened. The held-out material has a mechanism too: a paper composed as an assessment carries the role it plays, whether pre, post or retention, the questions drawn onto it leave the practice pool and nothing puts them back, and the sitting is graded without writing evidence, so it cannot move the mastery estimate it exists to check. A university's own mark can be recorded beside that result, and is admissible for analysis only where the consent row says so. The byline above is the ownership disclosure.
References
Butler, A. C. (2010). Repeated testing produces superior transfer of learning relative to repeated studying. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(5), 1118–1133. https://doi.org/10.1037/a0019902 Experiments.
Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267. Statistical methodology.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114 Research methodology perspective.
Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.
See the product a study would evaluate