Research
Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?
Overview Research area: Knowledge tracing (KT) with large language models — specifically the cold-start regime where a new course or platform has few or no logged learners. Technical level: Intermedia
- arXiv
- 2610.11135
- Published
- 2026-10-08
- Authors
- Unggi Lee, Haeun Park
AI summary
Overview
Research area: Knowledge tracing (KT) with large language models — specifically the cold-start regime where a new course or platform has few or no logged learners.
Technical level: Intermediate. The paper assumes familiarity with AUC, calibration metrics, and the difference between fine-tuned and prompted models, but its central argument is conceptual rather than mathematical.
Scope: A single study showing that an off-the-shelf "System-One" LLM (Jev), which returns a probability for a typed question in one pass without generating text, matches or beats supervised deep KT models and System-Two LLM-based KT baselines on seven public datasets when learner logs are scarce.
What This Paper Is About
Deep knowledge tracing models need many logged learners before they become useful, so a brand-new course or platform starts with no usable model. Most LLM-based KT approaches are "System-Two": the model either reasons, votes over ten sampled traces, or is fine-tuned on the target data, all of which are slow and produce only coarse probabilities. The authors ask whether a System-One LLM — one that directly returns a probability for a typed question in a single pass — can already perform KT when few or no learners have been logged.
Key Contributions
- Finding. With no learner logs at all, an off-the-shelf System-One LLM (Jev) reaches .706 mean AUC, above the best of 28 deep KT models trained on 8 learners (.689) and above the prompted System-Two method Thinking-KT (.650) on all seven datasets.
- Where the lead holds. A cold-start evaluation on identical targets across three settings (few training learners, new learners, new items) shows how long the advantage lasts as learners are logged, and introduces a training-free input recipe (JevKT) that keeps it until supervised KT takes over.
- Why the gain appears. Controlled analyses — reader swap, three other LLMs behind the byte-identical typed request, and contamination checks — indicate the gain comes from the model rather than from the input format or memorised data.
- Practical efficiency accounting. One JevKT prediction costs $886 per million predictions and one Jev 0-shot prediction costs $44, against $4,360 for Thinking-KT, with latency and throughput reported for every method.
Main Findings
- No learners logged: Jev 0-shot reaches .706 mean AUC with no data from the target platform, above the strongest training-free LLM-KT baseline LOKT (.671), above the best deep KT model trained on 8 learners (.689; higher on 5/7 datasets), and above Thinking-KT (.650) on all seven datasets.
- Cost of the zero-data result: Jev 0-shot uses one call instead of ten and costs about 1/100 of Thinking-KT's API cost.
- Advantage as learners are logged: JevKT (Jev plus 64 KC-matched examples plus a similar-learner kNN statistic) leads the best deep KT model by +.033, +.023 and +.016 mean AUC at N = 8, 16 and 32 (7/7 datasets in each case). The mean difference is +.004 at N = 64, −.002 at N = 128, −.008 at N = 256, and −.028 with full training data, so the crossover lies between 64 and 128 learners on average.
- Statistical significance: The lead is Holm-significant on 7/7 datasets at N = 8 and 6/7 at N = 16, falling to 3/7 at N = 32 and 2/7 at N = 64.
- New learners: JevKT is above the best deep KT model at every position in a new learner's sequence, including the second to fifth interactions (.714 vs. .691), and in 33/35 dataset–position cells.
- New items: When 20% of items are held out but all training learners are available, the best deep KT model is ahead (JevKT .713 and Jev 0-shot .704 vs. best DLKT .729; JevKT better on 1/7 datasets), while both still beat LOKT (.664). The authors conclude the advantage comes from having few learners, not from unseen items.
- Model, not format: On identical Qwen3.5-9B weights, a typed one-pass head (B0) is worse than a plain label-probability readout (A) — .630 vs. .688, and .701 vs. .752 with the similar-learner statistic. Distilling Jev into those weights (C, JEV-9B) does not beat A (.679 and .741; at N = 8, A .700, B0 .659, C .699). When Jev reads the identical serialised input, AUC rises in all 35 cells, for example from .727 to .776 with the JevKT input.
- Other readers fall short: Through the official System-One adapter, GPT-4o-mini (.680), Gemini-2.5-Flash-Lite (.670) and DeepSeek-V4-Flash (.626) are below Jev 0-shot (.706) on every dataset at N = 0. JEV-9B (.687 at N = 0, .699 at N = 8) and the open-source Laya (.536, near chance) are the only other System-One-interface readers tested.
- What the logged learners add: The input recipe helps Jev and DeepSeek by a similar small amount (+.016 and +.010 respectively), so the .077 gap at N = 0 is set by the model. At N = 8 the prior alone reaches .706, the similar-learner statistic alone .595, and both together .722. At N = 16 the statistic alone raises Jev 0-shot from .706 to .729, 64 KC-matched shots alone to .724, and both together to .730 — the statistic carries almost all of the gain.
- Reasoning does not help the backbone: With Qwen3.5-9B, a reasoning budget of 1024 or 2048 tokens never significantly improves AUC over the same model without reasoning (0/7 datasets), while one Jev pass beats Thinking-KT on all seven datasets with Holm-corrected ΔAUC of +.038 to +.073 against the 2048-token budget.
- Calibration: Thinking-KT's ten-vote probabilities pile up at 0 and 1 and have mean ECE .253 (LOKT .209), against .084 for a single Jev 0-shot pass. At N = 8, JevKT matches the best deep KT model (ECE .070 vs. .076; Brier .182 vs. .192) without post-hoc scaling. Jev's remaining error is a mild underconfidence.
- Longer reasoning marks hard cases: Within each dataset, the third of targets with the longest reasoning has lower AUC than the shortest third (.668 vs. .719 at budget 2048; .675 vs. .726 at 1024), and reasoning length correlates with absolute error (Spearman +.29).
- The gain is not confined to easy cases: At N = 8, JevKT leads on items the logged learners answered (+.036) and on items none of them answered (+.028), with the largest lead for learners whose accuracy so far is very low or very high (+.050 and +.063).
- Artefact checks: Relabelling item IDs and removing names changes Jev 0-shot AUC by at most .012 and is never significant; on six synthetic datasets generated after the model release, JevKT remains at or above the best deep KT model at N = 8; a leak audit passes in every setting, including a test that the input is unchanged when the target outcome is flipped.
- Baseline strength: Even with hyper-parameters tuned on validation learners (four strong pyKT models, eight configurations each), deep KT stays below JevKT by +.070, +.042, +.030 and +.010 at N = 8, 16, 32 and 64.
- Efficiency: One JevKT prediction takes 325 ms and costs $886 per million predictions; Jev 0-shot takes 200 ms and costs $44; LOKT costs $707; Thinking-KT needs $4,360 and 41.7 s per prediction. DLKT is far cheaper ($0.01 to $0.19 per million depending on model, with training on all learners costing at most $0.08) but must be retrained for every course and needs logged learners.
Methodology in Plain English
The authors treat KT as a single typed yes/no question posed to a frozen commercial model: "Will this learner answer the next item correctly?" A prompt is a JSON state containing the learner's last 25 interactions, the target item, one sentence with a similar-learner statistic, and 64 solved examples from logged learners whose items share a knowledge component with the target. The model's probability of "yes" is used directly, with no fine-tuning, no sampling and no post-hoc calibration.
They control what the question contains across two conditions: Jev 0-shot, which uses the learner's recent history alone and therefore no data from the target platform, and JevKT, which adds the 64 KC-matched examples and the similar-learner statistic. The statistic is the first-attempt accuracy on the target item of the 20 training learners whose smoothed per-KC accuracy vectors are closest in a weighted ℓ1 distance.
Evaluations run on seven public datasets with learners split 80/20 at the learner level. Three cold-start settings are scored on identical targets for all methods: P1 varies the number of training learners (N = 8, 16, 32, 64, 128, 256, all; 5 seeds; first 50 interactions of held-out learners, 2,000 targets per dataset); P2 re-slices the same N = 8 predictions by position in a new learner's sequence; P3 holds out 20% of items from training, examples and statistics (3 seeds, up to 1,000 targets). The deep KT comparison is the best of 28 pyKT models selected on test in every cell, which favours deep KT. Differences are tested with a one-sided paired learner-cluster bootstrap with Holm correction. The JevKT input configuration was selected on a development split by a rule fixed before test was evaluated.
To separate the model from the format, the authors hold Qwen3.5-9B weights fixed and compare a plain readout, a typed one-pass head, and a Jev-distilled student; they also swap in Jev as the reader on identical inputs, and send the byte-identical request to three other LLMs through the vendor's System-One adapter.
Why This Matters
Impact on research. The paper reframes the LLM-for-KT debate away from "reason more" or "fine-tune more" toward whether a model can simply read a probability from a typed question, and it reports calibration, latency and cost alongside AUC — quantities rarely reported together in LLM-KT work. It also supplies a conservative comparison protocol (best baseline chosen on test, validation-tuned controls, contamination and leak audits) for a claim that would otherwise look implausibly strong.
Real-world applications:
- Launching a new course, subject or item bank on a platform where no learner history yet exists.
- Serving next-item correctness predictions during the first sessions of a new cohort, before enough interactions accumulate for a supervised model.
- Deploying adaptive practice in low-resource or niche domains where thousands of learners per course are never available.
- Cross-platform transfer, where a model trained on one platform's logs does not fit a different item catalogue.
Industry relevance. The economics are the headline: JevKT costs $886 per million predictions against $4,360 for Thinking-KT, and Jev 0-shot costs $44, while both are ready at zero training time (0 / 0) for a new course. Supervised DLKT remains far cheaper per prediction ($0.01 to $0.19 per million) and faster (0.23 ms for DKT up to 3.08 ms for AKT), but its limit is data rather than money: it must be retrained per course and needs logged learners to train on. Under the pricing used here (one RTX 3090 at the RunPod Community Cloud on-demand rate of $0.22/h, September 2026), the trade-off is between a small per-call cost with instant availability and a near-zero per-prediction cost with a data and retraining prerequisite.
Future Directions
- Whether an open System-One model can match Jev remains open; Laya, the only open-source model tested that reproduces Jev's typed interface, stays near chance (.536 at N = 0, .553 at N = 8).
- Generality cannot be established across commercial System-One models because none other is available; the authors test generality only with other LLMs behind the official adapter, a Jev-distilled student and Laya, and note the adapter may disadvantage those readers because it reads probabilities written as text.
- Jev's training data are unknown, so memorisation cannot be ruled out despite the contamination checks; only one Jev version exists, so drift across model versions cannot be measured, though re-querying the pinned version after two days reproduces its predictions closely.
- The conclusions about reasoning apply to the Qwen3.5-9B backbone; a larger reasoning model may behave differently, and where exactly supervised KT takes over is hard to predict because a crossover model fitted on synthetic datasets does not transfer to the real ones.
Target Audience
Researchers and engineers working on knowledge tracing, educational data mining, or LLM-based prediction under cold start, and practitioners deciding whether to deploy a per-call LLM prediction service or invest in per-course supervised training. It also suits evaluation-focused readers interested in how to test whether a strong model result comes from the model or from the prompt format.
Authors’ abstract
Knowledge tracing (KT) models need many logged learners, so a new course or platform starts without a usable model. In LLM-based KT the LLM generates the answer, which we call System-Two; it is either fine-tuned on the target data or reasons and votes over ten samples, which is slow and gives coarse probabilities. We ask whether an off-the-shelf System-One LLM, which returns a probability for a typed question directly in a single pass, can perform KT when few or no learners are logged. On seven datasets, Jev without any data from the target platform reaches a mean AUC of .706, above the best of 28 deep KT models trained on 8 learners (.689) and above System-Two Thinking-KT on all seven datasets (.650) at about 1/100 of its API cost. Adding examples and a similar-learner statistic from the logged learners (JevKT) raises this to .722; JevKT stays significantly ahead of deep KT up to 16 learners and ahead on average up to 64, and supervised KT catches up between 64 and 128 learners. Among the readers we tested, the gain is specific to Jev, since three other LLMs queried with the byte-identical typed request through the official System-One adapter fall below it on all seven datasets, and reader swaps and contamination checks find no evidence that the input format or memorised data explain the gain. For new learners the advantage holds from their first interactions, whereas on unseen items with all learners logged, deep KT remains ahead.