2026-08-06
How Multi-Concept Sentence Reviews Should Be Scheduled
Schedule multi-concept sentence reviews by packing due words into novel sentences and grading each concept separately so FSRS advances only what you know.
The short answer
Multi-concept sentence reviews should be scheduled at the concept level, not the sentence level. Pack as many due words and grammar items as you can into one natural, novel sentence; grade each item separately; then let a scheduler such as FSRS update only the components that were actually retrieved or missed. The sentence is the exercise shell. The due list and the per-item grades are what make the schedule honest.
That design is often called concept stacking: one production attempt clears several reviews at once without turning the whole sentence into a single pass/fail card. LinGoat builds reviews this way: native-to-target translation prompts that pack due concepts into fresh sentences, plus word-by-word grading that feeds FSRS. It is written sentence practice with spaced repetition, not a speaking replacement.
What concept stacking is
Concept stacking means the review engine looks at which vocabulary and grammar items are due today, then builds a sentence that includes as many of those items as possible while staying natural. Instead of five isolated flashcards for five words, one well-built sentence can exercise a verb form, a preposition, and several lemmas in a single retrieval attempt.
The goal is efficiency without fake fluency. You still practice real composition: choosing forms, agreement, and word order under a production prompt. You just amortize that effort across more due items per minute. For a hands-on workflow (prompt, write, grade, reschedule), see how to practice sentence construction with spaced repetition.
Why sentences beat isolated word reviews
Words do not live in a vacuum. Their meanings, collocations, and nuances shift with surrounding syntax. Nation argues that vocabulary knowledge is multi-faceted: deliberate study can build form-meaning links efficiently, while sentence-level encounters build usage knowledge that bare translation pairs rarely provide.1
Stacking due items inside sentences keeps that contextual benefit while attacking the review queue. You are not choosing between "context" and "spacing." You are using context as the delivery format for spaced components. Related evidence on dynamic sentence packing (with independent per-word scheduling) is summarized in our sentence-based spaced repetition research post; this article focuses on the scheduling logic itself: how stacking, grading, and FSRS fit together.
Novel sentences, not mined static ones
Stacking only works if the sentence changes. If you always review the same mined example, you can pass by recognizing the string rather than retrieving the components. Encoding specificity research shows that retrieval often depends on cues present at encoding; when those cues disappear, recall can fail even though the "card" felt easy.2
Transfer-appropriate processing makes the same point from the opposite direction: practice that matches the later task (building new sentences) transfers better than practice that only matches a fixed flashcard frame.3 So the stacked review should be a novel sentence: same due concepts, new syntax and collocations. Why static mining breaks SRS is covered in more depth in why sentence mining fails.
Why packing requires granular grading
Here is the failure mode that kills stacked reviews. Suppose a sentence packs five due concepts and you miss one article. If the whole sentence is marked "Again," four solid recalls get punished. If the whole sentence is marked "Good," the miss never returns soon enough. Either way, the scheduler receives noise.
Concept stacking therefore requires granular attribution: each word and grammar concept in the answer is graded on its own, then scheduled on its own. Only then can packing stay efficient. The sentence can be dense; the memory model stays atomic. For the grading argument in isolation, see word-by-word grading in language learning.
FSRS schedules components, not sentences
Once grades are per concept, the timing engine can do its job. FSRS (Free Spaced Repetition Scheduler) models each item with difficulty, stability, and retrievability rather than a single SM-2 ease factor.4 On large Anki-collection benchmarks, FSRS predicts recall more accurately than SM-2, which lets intervals sit closer to the forgetting curve without over-reviewing.5
In a stacked sentence session, FSRS does not schedule "Sentence #482." It schedules the lemmas and grammar rules inside it. Items you retrieve correctly drift out; items you miss return sooner and can be re-packed into a different novel sentence next time. That is how stacking and scheduling reinforce each other instead of fighting each other.
Constraints: naturalness and cognitive load
Stacking has limits. A sentence stuffed with every due item can become unnatural or overload working memory. Cognitive load research warns that tasks with many interacting elements (new forms plus full syntax plus production) can exceed capacity and produce frustration rather than learning.6
Good multi-concept reviews therefore optimize under two constraints:
- Natural phrasing: pack due concepts only when they fit a plausible sentence a learner might actually say or write.
- Manageable interactivity: prefer several dense-but-readable sentences over one impossible megasentence, especially for newer items still ramping into full sentence production.
Efficiency is "more due concepts cleared per honest retrieval," not "maximum tokens per prompt."
How LinGoat schedules multi-concept reviews
LinGoat generates native-language prompts that pack currently due vocabulary and grammar into novel target-language sentences. You produce the full sentence; each concept is graded individually; FSRS updates those components. The next session rebuilds fresh sentences from whatever is due again. For the full three-part loop (spaced repetition, dynamic exercises, granular attribution), see LinGoat's full pedagogy.
See how LinGoat works or try the app.
References
- Nation, I. S. P. (2001). Learning Vocabulary in Another Language. Cambridge University Press. https://doi.org/10.1017/CBO9781139524759
- Tulving, E., & Thomson, D. M. (1973). Encoding specificity and retrieval processes in episodic memory. Psychological Review, 80(5), 352-373. https://doi.org/10.1037/h0027317
- Morris, C. D., Bransford, J. D., & Franks, J. J. (1977). Levels of processing versus transfer appropriate processing. Journal of Verbal Learning and Verbal Behavior, 16(5), 519-533. https://doi.org/10.1016/S0022-5371(77)80016-9
- Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 4384-4394). https://doi.org/10.1145/3534678.3539081
- Expertium. (2025). Benchmark of spaced repetition algorithms. https://expertium.github.io/Benchmark.html
- Sweller, J. (2010). Element interactivity and intrinsic, extraneous, and germane cognitive load. Educational Psychology Review, 22(2), 123-138. https://doi.org/10.1007/s10648-010-9128-5