Back to blog

2026-08-05

Word-by-Word Grading: Why Per-Concept Feedback Matters for SRS

Word-by-word grading scores each concept so SRS can reschedule failures. Whole-sentence pass/fail contaminates FSRS and wastes review time on known items.

The short answer

Word-by-word grading matters for spaced repetition because SRS algorithms schedule items, not whole sentences. If you mark an entire sentence wrong when only one word failed, every other concept in that sentence gets punished. If you mark it right when one grammar form was off, the weak item gets a free pass. Either way, the schedule drifts away from what you actually know.

Per-concept feedback (what LinGoat calls granular attribution) grades each vocabulary and grammar concept in the sentence on its own, then feeds those outcomes into the scheduler. That is how multi-word sentence practice can stay compatible with FSRS-style timing instead of fighting it. For the broader loop, see LinGoat's full pedagogy and how spaced repetition works.

What word-by-word grading is (granular attribution)

In a translation exercise you produce a full target-language sentence. Word-by-word grading does not stop at "correct" or "incorrect" for the string as a whole. It attributes success or failure to each tracked concept inside that string: a verb form, a preposition, a noun, a tense marker, and so on.

Granular attribution is LinGoat's name for that step. The learner-facing idea is simple: you should know which piece failed, and the app should reschedule that piece. Immediate, localized correction also helps you update the hypothesis you just tested, rather than leaving you with a vague sense that "the sentence was wrong."

Why whole-sentence pass/fail breaks SRS

Spaced repetition assumes each review event is a clean signal about one memory. Modern schedulers such as FSRS model retrievability, stability, and difficulty per item and update those variables from your grade on that item.1 A sentence that packs five due concepts into one exercise is efficient for practice, but only if the system can split the outcome five ways.

Whole-sentence pass/fail collapses those five signals into one bit:

  • False failures: You miss one ending and four solid words get treated as forgotten. Intervals shrink for material you already retrieve well. Review load rises for the wrong reasons.
  • False successes: You nail most of the sentence and the one shaky concept rides along as "Good." The scheduler lengthens that item's interval. You see it again too late.
  • Unusable averages: Even a middle score ("Hard" for the whole sentence) does not tell FSRS which concept was hard. Difficulty and stability updates smear across unrelated memories.

That is not a minor UX preference. It is a data-integrity problem. Adaptive apps that claim to track mastery need item-level outcomes; a sentence-level thumbs-up cannot substitute. See what adaptive language apps should track.

What FSRS needs from each review

FSRS (and related spaced-repetition models) are only as good as the review log they receive. Ye, Su, and Cao (2022) frame scheduling as optimizing when to ask again so that predicted recall stays near a target retention level.1 That optimization assumes each logged grade corresponds to a specific memory unit.

Benchmarks that compare FSRS with older SM-2-style heuristics likewise evaluate prediction error on card-level review histories, not on unlabeled paragraph scores.2 If your "cards" are really multi-concept sentences and you only store pass/fail for the blob, you are asking a per-item model to train on corrupted labels.

Retrieval practice research points the same direction from the learning side: the test event that strengthens memory is a retrieval attempt for the information you care about, not a global vibe check on a paragraph.3 Sentence production can still be the exercise format. The grading layer has to resolve which concepts were retrieved successfully.

Recasts at adult scale (optional analogy)

In first-language development, caregivers often reformulate a child's error with a localized correction (a recast): the child says "I eated," and the adult replies "you ate." Research on adult reformulations treats that kind of feedback as negative evidence that helps revise structural hypotheses.4

Adult learners rarely get thousands of spoken recasts per day. Word-by-word grading on written production is not identical to caregiver talk, but it plays a related role: it localizes the failure so you can update the right rule instead of discarding the whole attempt. LinGoat uses that signal for scheduling as well as for feedback. It is written sentence practice plus SRS, not a claim that the app replaces conversation partners.

How this fits translation and multi-concept reviews

Granular attribution is why native-to-target translation exercises can feed a scheduler at all. Translation forces you to attempt the due concepts instead of avoiding them; word-by-word grading turns that forced attempt into clean per-concept outcomes.

The same logic unlocks multi-concept sentence reviews: packing several due items into one natural sentence saves time only when each item can succeed or fail independently. Without that split, "efficient" stacking becomes efficient contamination.

LinGoat generates novel sentences around your due concepts, grades each concept individually, and schedules with FSRS. That combination is the core loop: spacing, production, and granular attribution together. See how LinGoat works or try the app.

References

  1. Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 4384-4394). Association for Computing Machinery. https://doi.org/10.1145/3534678.3539081
  2. Expertium. (2025). Benchmark of spaced repetition algorithms. https://expertium.github.io/Benchmark.html
  3. Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
  4. Chouinard, M. M., & Clark, E. V. (2003). Adult reformulations of child errors as negative evidence. Journal of Child Language, 30(3), 637-669. https://doi.org/10.1017/S0305000903005701