memoza

The Memoza AcademyEngineering

Why a PDF is only the starting point for Memoza

A hundred fluent questions is not a question bank. Why generation needs a specification, and what review has to ask of each item.

By Lucas Hsu · · 5 min read


Imagine uploading lecture slides on regression and receiving a hundred questions. Most use the right vocabulary. Some have clear answers. A few ask about the incidental example used to introduce a concept.

The list looks substantial. Its educational value remains uncertain.

For an exam-preparation product, question quality depends on more than a connection to the source text. We need to know which capability a question exercises, whether the response reveals that capability, and how the activity fits the course. This is why the architecture we propose for Memoza treats document ingestion as the beginning of question design.

Plausibility is an incomplete specification

A slide can contain a definition, a worked example, an illustrative anecdote, and an administrative reminder. Each is meaningful in context. They do not deserve equal representation in a practice bank.

Even an accurate question can target the wrong performance. A lecture about regression assumptions might lead to “List the assumptions.” The exam may instead present a study and ask which assumption is questionable and how that affects interpretation.

The first question has a legitimate use. The problem arises if success on it is treated as evidence that the second task is covered.

Direct generation should not be dismissed categorically. Säuberli and Clematide found promising results from language models generating reading-comprehension questions, while evaluating qualities such as answerability and whether the source text was actually needed (Säuberli and Clematide, 2024). Their findings concern that task and evaluation setting. They support investigating generated questions carefully, rather than declaring either that simple prompting always fails or that fluent output is sufficient.

A specification makes errors easier to detect

Automatic item generation has a history before generative AI. Gierl, Lai and Turner described a process that begins with a cognitive model, develops item models, and then uses software to generate test items (Gierl et al., 2012). Their work in medical assessment illustrates the value of structuring the intended task before producing variants.

For a regression activity, a useful specification might require the student to identify a violation of independence in a repeated-measures setting and explain the implication for a standard analysis. It could also define the expected response length, acceptable reasoning, and misconceptions that should not receive full credit.

That specification makes a generated question reviewable. A reviewer can inspect whether the scenario contains enough information, whether a competing interpretation is equally reasonable, and whether the marking criteria match what was asked.

Changing the names or numbers in a question does not necessarily change its reasoning demands. Conversely, a small change to the assumptions can turn a routine calculation into a different problem. Variation should therefore be checked at the level of the task, as well as the wording.

Changing the names or numbers in a question does not necessarily change its reasoning demands.

Keep coverage and priority distinct

The course model should answer what belongs to the syllabus. A study policy should decide what currently deserves attention. An assessment profile should describe the evidence available about exam demands.

These decisions interact, but they should remain inspectable. Otherwise, a topic that appears rarely in the source files may disappear from practice, or an old paper may become a misleading prediction of the next exam.

The same separation helps with source use. Practice should develop the intended skills through genuinely authored tasks. Provenance and similarity checks can help reviewers identify copied material, but neither a low similarity score nor a new set of numbers establishes permission to reuse a source. Rights decisions require their own review.

Quality assurance needs several questions

For Memoza's proposed workflow, the first check is substantive correctness: can the question be answered, and does the answer follow from the stated conditions? The next is alignment: does answering it exercise the intended objective? Then comes grading: would a valid alternative solution receive appropriate credit?

Difficulty requires another level of evidence. An author's estimate is useful for initial placement. Observed difficulty depends on the learners, their preparation, the available support, and the other items around it. A question that is hard because of ambiguous wording should not be celebrated as advanced.

Finally, the bank needs review as a whole. A hundred individually acceptable items can still leave a major objective uncovered. They can also repeat the same easy procedure while giving the impression of variety.

The proposed standard is therefore traceability from source and objective to task and response. A reviewer should be able to explain why an item belongs, what success would show, and what it leaves untested.

Question generation becomes more useful when those decisions constrain it. The size of the bank is then a consequence of a practice plan, rather than the main evidence that a course has been understood.

How Memoza fits

In the course builder, the part that decides which questions should exist writes none of them: it is a deterministic planner that produces specifications, each naming the concept the question must assess, the learning objectives behind it, the response format, the intended difficulty and the review it will need. It invents nothing to fill a gap, so a course whose exam profile states no response format is blocked rather than given a default, and one that states no topic weights gets even coverage with the plan saying so. One provider call then drafts one item for one specification, and a draft that cannot be assembled into a valid question is refused whole rather than patched; where a draft proposes the values behind its rubric steps, a deterministic engine parses and verifies each one before anything is stored. Review is a list of named checks rather than a single score, a check that could not be run is recorded as unknown and never as a pass, and the course reaches a student only after nine named gates pass, one of which asks that every planned question exists, is current and carries clean evidence.

References

Gierl, M. J., Lai, H., & Turner, S. R. (2012). Using automatic item generation to create multiple-choice test items. Medical Education, 46(8), 757–765. Assessment generation methodology.

Säuberli, A., & Clematide, S. (2024). Automatic generation and evaluation of reading comprehension test items with large language models. READI workshop at LREC-COLING 2024. Author manuscript: arXiv:2404.07720, version 2. Conference research paper.

Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.

See the courses this process produces
  • Engineering · 5 min read

    Building a compiler for university courses

    Slides and past papers are not a learning environment. The reviewable course representation we build before any practice is planned.

    Read guide →

  • Engineering · 5 min read

    Choosing the next question in Memoza

    Eligibility first, then the need behind the activity, then difficulty that serves the purpose. A recommendation you can reconstruct.

    Read guide →

Put the ideas to work.

Explore course practice in Memoza.

Try the demo, up to 3 questions

Explore course practice