Research
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
Overview Research area: Natural Language Processing / natural language generation applied to computer-assisted language learning — specifically constrained text generation with large language models.

- arXiv
- 2512.18362
- Published
- 2025-12-20
- Authors
- Wiktor Kamzela, Mateusz Lango, Ondrej Dusek
AI summary
Overview
Research area: Natural Language Processing / natural language generation applied to computer-assisted language learning — specifically constrained text generation with large language models.
Technical level: Intermediate. The paper assumes familiarity with prompting strategies, lexical constraints, and NLG evaluation, but the core ideas (limiting a story to known vocabulary, repeating target words) are explained in plain terms. The appendices include full prompts and metric definitions.
Scope in one sentence: The paper introduces SRS-Stories, a system in which an LLM generates graded-reader-style stories constrained to a learner's known vocabulary plus words selected by a Spaced Repetition System, and evaluates three story-generation strategies and three lexical-constraint-enforcement strategies in English, Chinese, and Polish.
What This Paper Is About
Spaced Repetition Systems (SRS) such as Anki, SuperMemo, or HackChinese are effective for recalling flashcards but weak at teaching real vocabulary acquisition, because words are reviewed out of context and the process becomes tedious. Graded readers solve the context problem, but finding one at the right level is hard (especially outside English) and their repetition of words is not optimized.
The paper's goal is to generate graded readers automatically: given the words a learner already knows (V) and the words an SRS has selected to learn (L), produce a coherent, enjoyable story that uses only V ∪ L and includes each new word at least c times (c = 3 in the paper), so reading the story doubles as vocabulary teaching and review.
Key Contributions
-
Three prompting strategies for story generation — Simple Prompting, Planning, and Examples First — designed to produce coherent stories that teach words arbitrarily selected by an SRS, with each target word appearing at least three times in meaningful context.
-
Three strategies for enforcing vocabulary constraints based on text rewriting, using an external non-neural constraint verifier and iterative re-generation: Rewrite (base), Rewrite Highlighted, and Get Synonyms then Rewrite.
-
An experimental evaluation across three languages (English, Chinese, and Polish) at multiple language proficiency levels, comparing the proposed methods against Constrained Beam Search (CBS) as a baseline.
-
Human evaluation alongside LLM-based evaluation, comparing Simple Prompting to CBS on 50 English, 50 Chinese, and 50 Polish stories, and reporting correlations between human and LLM judgments.
Main Findings
-
CBS is the weakest method. Constrained Beam Search scored lowest on grammaticality (3.22), coherence (3.82), and interestingness (3.92) among the English B1 methods in Table 1. Manual analysis found CBS stories sometimes list new words out of context, are not fluent, contain frequent grammatical errors, and change plots unexpectedly (e.g., suddenly changing the protagonist).
-
Simple Prompting was strong across quality metrics. At English B1 it achieved the highest coherence (4.22), second-best grammaticality (4.05), and highest interestingness (4.01, tied with Planning). Averaged over B1/B2/C1 it scored 4.04 grammaticality, 4.27 coherence, 4.06 interestingness.
-
Planning used target words most often. At English B1, Planning reached an average of 2.68 occurrences per word to learn (versus 1.92 for Simple Prompting and 1.46 for CBS) and produced longer stories (609.67 words), but with lower grammatical accuracy (3.98).
-
Examples First maximized grammaticality but minimized exposure. It had the highest grammar score at B1 (4.08) but the lowest coherence and interestingness, and the lowest occurrence count (1.42), exposing the learner least frequently to target vocabulary.
-
The basic Rewrite strategy gave the best constraint satisfaction. With base Rewrite, all generation methods reached their best grammaticality, highest #L, and lowest OOV percentages. The more advanced rewriting approaches produced more interesting and sometimes more coherent stories that were longer but contained more words unfamiliar to the user.
-
Statistically significant improvement over CBS. Using two-sample unpaired t-tests on English results averaged across levels, differences between CBS and any proposed strategy were significant across all three LLM-measured quality aspects, with p < 0.001. Differences among lexical constraint enforcement strategies were not statistically significant for grammaticality.
-
Multilingual results vary by language. For Chinese, Simple Prompting with base Rewrite had the highest grammaticality (3.99) while Rewrite Highlighted gave higher coherence (4.09), interestingness (3.73), and the most frequent inclusion of new words (3.13 occurrences). For Polish, Rewrite Highlighted achieved the highest grammaticality (3.81) and coherence (4.08) but a higher OOV rate (3.45%) than base Rewrite (0.34%).
-
Llama 3.1 70B Instruct was best at vocabulary control. Among the compared models it was the most successful at eliminating unknown words (OOV 0.52% with Simple Prompting) and at including all new vocabulary (95.22% of words to learn introduced).
-
Model comparison caveats. Qwen 2.5 scored high on LLM-based metrics (e.g., coherence 4.76 with Simple Prompting) but may be biased since Qwen was also the evaluator. Aya 23 35B achieved the highest grammaticality, coherence, and interestingness scores (up to 4.21/4.84/4.49 with Rewrite Highlighted) but was ineffective at including new words (33.52% with Simple Prompting). DeepSeek R1 70B did best with base Rewrite, and more advanced rewriting harmed its performance.
-
Human evaluation favored the proposed method on the two learning-critical aspects. Against CBS on English, stories from Simple Prompting were significantly more grammatically correct (p = 0.0151) and significantly better at illustrating the use of studied words (p < 0.0001). Human scores: English CBS 3.56 grammar, 3.96 coherence, 3.23 interestingness, 1.24 word use, 3.24 overall; English Ours 4.08, 3.92, 2.94, 3.88, 3.08. Note the interestingness and overall scores were not higher for the proposed method in this comparison.
-
Chinese stories were rated highest by humans. Chinese (Ours) received 4.68 grammar, 4.40 coherence, 3.66 interestingness, 4.40 word use, 4.16 overall. Polish (Ours) received 3.78, 3.92, 3.50, 3.80, 3.66.
-
LLM-based metrics correlate only moderately with humans. Pearson correlations for English were 0.557 (grammar), 0.507 (coherence), 0.456 (interestingness); Spearman correlations were 0.383, 0.280, 0.465. Correlations were lower for Chinese and Polish (Polish interestingness Pearson was -0.149).
Methodology in Plain English
The researchers built a pipeline around an LLM that writes a story from a list of target words, then checks and repairs the vocabulary. The check is done not by the LLM but by a separate verifier that normalizes text, tokenizes, and lemmatizes it, then compares the words used against the allowed set (known words plus selected new words). Any word outside that set is flagged as a broken constraint.
Three prompting approaches were tested. Simple Prompting asks directly for a 500–750 word story using each target word at least three times in context. Planning splits the job into generating story title ideas, picking the most interesting and writing an outline, then turning the outline into the final story. Examples First asks the model to write an example sentence for each target word before writing the story.
Three repair strategies were tested. Rewrite simply lists the unknown words and asks the model to replace them with simpler alternatives, repeated five times (the authors note replacements flatten out after 5 iterations). Rewrite Highlighted shows the story with unknown words wrapped in asterisks (Markdown bold) and asks the model to simplify those. Get Synonyms then Rewrite first asks for synonyms of the unknown words, then asks for a rewrite using them.
Vocabulary levels came from CEFR-J word lists for English, New HSK 3.0 lists for Chinese, and, because no official Polish lists exist, frequency lists derived from the Polish Wikipedia portion of the Leipzig Corpora Collection (top 1,500 words = A1, 3,500 = A2, 5,000 = B1, 7,500 = B2, 10,000 = C1, remainder = C2). Each simulated study session assumes the learner knows everything at a given level and wants to learn 10 random words from the next level; results are averaged over 200 stories per language level and method.
Stories were generated with Llama 3.1 70B Instruct using vLLM default parameters except a maximum of 4096 tokens and temperature set to zero. Qwen2.5 72B Instruct, Aya 23 35B, and DeepSeek R1 70B were also compared. The baseline was the HuggingFace implementation of Constrained Beam Search combined with the "No Bad Words" token-masking strategy.
Evaluation used Qwen2.5 72B Instruct as an LLM judge scoring grammatical correctness, coherence, and interestingness from 1 to 5 after first listing errors, plus count-based metrics: average occurrences per target word (#L), percentage of target words actually introduced (#|L| ≥ 1), story length, and OOV percentage. Appendix D adds new-word-to-length ratio, perplexity (from Qwen2.5 7B), a BERT model finetuned on CoLA for sentence acceptability, and BERT's next-sentence-prediction head for coherence — the latter three computed for English only. The human study used 20 annotators recruited via Prolific who had studied the story language as a second language and reached fluency; each evaluated 10 stories plus one attention check, and annotations were estimated at 45 minutes for eleven stories.
Why This Matters
Impact on research. The paper shows that prompt-based LLM generation with an external, non-neural constraint verifier can outperform a classic constrained decoding method (CBS) on grammar, coherence, and word-usage exemplification for lexical-constraint story generation. It also contributes a multilingual evaluation setup (English, Chinese, Polish) and an explicit check of how well an LLM judge agrees with human raters in a multilingual setting, where agreement was only moderate to low.
Real-world applications:
- Replacing or supplementing flashcard review in SRS apps, where daily reviews become short personalized stories rather than isolated word prompts.
- Producing graded readers on demand, particularly for languages where suitable readers are scarce (the paper explicitly notes this difficulty for languages other than English).
- Generating targeted practice material that reviews words an SRS predicts are at risk of being forgotten.
- Providing vocabulary-controlled reading material for learners at specific CEFR or HSK levels, with context-based exposure to function words that are hard to learn from flashcards.
Industry relevance. The system is designed to couple directly to existing SRS engines (Anki, SuperMemo, HackChinese are named), meaning an LLM story generator could be layered onto an installed vocabulary database. The comparison of Llama 3.1 70B, Qwen2.5 72B, Aya 23 35B, and DeepSeek R1 70B gives practical guidance on model choice, and the finding that simple base Rewrite often suffices suggests cheaper pipelines. The paper notes that more advanced generation strategies cost more LLM iterations, which is a deployment consideration.
Future Directions
- Reducing out-of-vocabulary words further, especially at lower proficiency levels, where the paper states generation is more challenging due to stronger lexical constraints, and in Polish and Chinese where OOV rates were noticeably higher than in English.
- Improving bias mitigation in generated content; the limitations section notes stories may reflect social biases from pretraining data and may contain minor inconsistencies or unnatural phrasing.
- Extending the system to phrases and idioms, which the FAQ says requires no technical modification, and reconsidering mixed new-plus-review sessions, which preliminary experiments found much harder than teaching new words alone.
- Improving multilingual LLM-as-judge reliability, given the low human–LLM correlations for Chinese and Polish and the potential self-preference bias when Qwen was both generator and evaluator.
- Validating the approach with actual learners over time, since the experiments simulate a student who knows all words at a level rather than using real SRS histories; the FAQ acknowledges this gap between assumed and actual known vocabulary.
Target Audience
Researchers and practitioners in natural language generation, computer-assisted language learning, and educational technology who are interested in lexically constrained text generation, LLM prompting strategies, or SRS integration. It is also relevant to developers building language-learning applications who need to choose between constrained decoding and prompt-based generation, and to researchers studying the reliability of LLM-based evaluation across multiple languages. Readers without NLP background can follow the problem framing and results but will need some familiarity with metrics and prompting terminology for the method and appendix details.
Authors’ abstract
In this paper, we use large language models to generate personalized stories for language learners, using only the vocabulary they know. The generated texts are specifically written to teach the user new vocabulary by simply reading stories where it appears in context, while at the same time seamlessly reviewing recently learned vocabulary. The generated stories are enjoyable to read and the vocabulary reviewing/learning is optimized by a Spaced Repetition System. The experiments are conducted in three languages: English, Chinese and Polish, evaluating three story generation methods and three strategies for enforcing lexical constraints. The results show that the generated stories are more grammatical, coherent, and provide better examples of word usage than texts generated by the standard constrained beam search approach