Research
ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI Tutors
ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI Tutors Overview Research area: AI in education / educational natural language processing — specifically, building training

- arXiv
- 2601.08950
- Published
- 2026-01-13
- Authors
- Mayank Sharma, Roy Pea, Hari Subramonyam
AI summary
ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI TutorsOverview
Research area: AI in education / educational natural language processing — specifically, building training data for dialogic tutoring behavior in large language models.
Technical level: Intermediate. The paper is readable without deep learning-sciences background, but understanding the evaluation requires familiarity with fine-tuning (QLoRA), regression classifiers, Pearson correlations, and linear mixed-effects models.
Scope: The paper releases ConvoLearn, a dataset of 2,134 semi-synthetic tutor–student dialogues labeled across six theory-grounded dialogic tutoring dimensions, validates that its labels carry pedagogical signal against authentic classroom transcripts, and demonstrates that fine-tuning a 7B open-weight model on a high-quality subset (1,250 dialogues) produces tutoring that credentialed teachers rate as comparable to a strong proprietary baseline.
What This Paper Is About
LLMs are increasingly used as tutors, but because they are trained to be broadly "helpful" assistants, they tend to explain answers directly rather than draw out student thinking — the opposite of what effective tutoring research says works. Existing responses to this problem are either prompting (which changes behavior only at inference time), evaluation benchmarks (which diagnose but do not correct), or proprietary fine-tuning data (which is not publicly released). ConvoLearn fills the missing middle: an openly available, dimension-labeled dialogic tutoring dataset grounded in knowledge-building theory, built from California middle school Earth Science content, with per-dialogue quality ratings preserved.
Key Contributions
-
A public dataset: ConvoLearn contains 2,134 semi-synthetic tutor–student dialogues (~20 turns each), with teacher turns authored by credentialed human teachers and student turns simulated, labeled across six dimensions and 21 subdimensions of dialogic tutoring. Quality ratings are preserved across the full spectrum to support fine-tuning, classifier training, and contrastive learning.
-
Ecological validity evidence: A Longformer classifier trained on ConvoLearn and applied to authentic K-12 classroom transcripts from the NCTE corpus produces scores that correlate significantly with expert-coded instructional quality on five of six MQI subscales and two of five CLASS subscales (range r = .118–.258, all p < .05), even though the classroom data is mathematics instruction and ConvoLearn is Earth Science tutoring.
-
A fine-tuning proof of concept: Open-weight 7B–8B models (Qwen-2.5-7B-Instruct, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3) were trained with multi-turn QLoRA on the high-quality subset. Credentialed teacher raters judged the fine-tuned Mistral-7B as comparable to Claude Sonnet 4.5 on a 1–5 effectiveness scale.
-
A replicable collection pipeline: The paper documents a scalable process — 323 credentialed K-12 teachers with a mean of 10.9 years of experience, recruited via Prolific with qualification quizzes and multi-stage quality filtering — showing dimension-labeled dialogic data can be collected without prohibitive cost or expertise barriers.
Main Findings
-
LLMs remain pedagogically misaligned: The paper cites prior work showing models score high on surface qualities like coherence and human-likeness while underperforming human tutors on dialogic dimensions such as providing meaningful guidance and engaging students in genuine reasoning.
-
Data scale and coverage: Of 500 recruited teachers, 72 failed the qualification quiz and 323 completed all tasks. From 3,076 extracted conversations, filtering (complete dialogues, keystroke-based AI-generation flag at a threshold of >0.5, duplicate and error removal) yielded 2,155 conversations, then 2,134 after safety filtering and 1,518 after quality filtering. Dimension coverage in the full dataset is uneven: metacognition 27.6%, cognitive engagement 23.5%, formative assessment 13.6%, power dynamics 13.4%, accountability 12.9%, cultural responsiveness 9.0%. Cultural responsiveness was the hardest to operationalize because teachers struggled to connect content to a simulated student of unknown background.
-
Annotation reliability and two dataset configurations: Dual LLM annotation (GPT-4o and Claude Haiku) reached moderate agreement (quadratic κ = 0.68, Spearman ρ = 0.72). Safety filtering removed 21 conversations; quality filtering removed 616. Claude Sonnet 4.5 adjudicated larger rating discrepancies, invoked in 10.2% of effectiveness and 2.9% of completeness cases. The full dataset has mean Effectiveness 3.6 and 71.4% fully complete conversations; the high-quality subset of 1,250 has mean Effectiveness 4.10/5 and 82.8% fully complete.
-
Classifier performance on held-out data: The Longformer regression model (initialized from allenai/longformer-base-4096, 5 epochs, dialogues partitioned 70/15/15 by seed question to prevent leakage) achieved Pearson r = 0.736 (p < .001), RMSE = 0.710, and MAE = 0.530 on its held-out test set.
-
Cross-domain validation signal: Applying that classifier to NCTE transcripts segmented into 20-turn chunks and averaged to the teacher-year level (n = 322), the strongest associations were with MQI's ETCA (r_CE = .258, r_FA = .250, r_MC = .224, all p < .001), SMQR (r_CE = .243, r_FA = .231, r_MC = .199, all p < .001), and EXPL (r_CE = .219, r_FA = .213, r_MC = .194, all p < .001). Smaller associations held for MLANG and LINK. MGEN was not significant across any dimension (r = .048–.079), consistent with its known label sparsity in NCTE.
-
Metacognition is harder to detect in whole-class data: Among CLASS subscales, CLAPS and CLQF reached significance for cognitive engagement and formative assessment (r_CE = .130, r_FA = .125, both p < .05) but not metacognition. CLCU, CLSTENG, and CLINSTD did not reach significance.
-
Fine-tuned 7B model is competitive with a proprietary baseline: With 31 certified teacher raters evaluating four-turn dialogues across 72 unseen seed questions (42 Earth Science, 30 Physics), fine-tuned Mistral-7B scored M = 3.49 versus Claude Sonnet 4.5 at M = 3.56 (β = −0.07, SE = 0.12, z = −0.55, p = .583) — a non-significant difference.
-
Gemini 2.0 Flash scored significantly higher: M = 3.82 (β = 0.27, SE = 0.12, z = 2.25, p = .025 versus Mistral), which the authors attribute plausibly to LearnLM's pedagogy-informed post-training at scale rather than imitation fine-tuning on 1,250 dialogues.
-
Simulator sensitivity check: The fine-tuned model was evaluated with GPT-4o as an alternative student simulator across three pedagogically distinct student profiles (engaged, limited prior knowledge, disengaged), with performance above M > 4.0 across all three.
Methodology in Plain English
The researchers first turned 60 multiple-choice questions from a freely available California Earth Science Standards Test into first-person student "doubts" to start conversations, covering Investigation and Experimentation, Astronomy and Cosmology, Solid Earth, and Earth's Energy Systems. Credentialed teachers then tutored a simulated student named "Jamie," powered by Gemini-1.5-Pro with a fixed system prompt, through a custom web platform; each teacher was randomly assigned two of the 21 subdimensions, passed a qualification quiz, and completed six conversations of 20 turns each (10 teacher, 10 student).
Conversations were cleaned in stages: incomplete dialogues under 10 teacher turns were dropped, keystroke logging flagged likely AI-generated responses, and duplicates and errors were removed. Each remaining dialogue was then independently scored by GPT-4o and Claude Haiku against a structured rubric covering effectiveness, completeness, quality issues, and safety issues. Harmful dialogues were removed, dialogues both annotators flagged for the same quality problem were excluded, and rating discrepancies of more than one point were adjudicated by Claude Sonnet 4.5. The authors describe the resulting labels as "silver-standard" annotations rather than gold-standard human labels.
To test whether the dataset captured real pedagogical substance rather than surface patterns, the team trained a Longformer regression model on ConvoLearn to predict effectiveness scores, then ran it over authentic NCTE classroom transcripts and correlated the resulting teacher-year scores with two established observation instruments, MQI and CLASS. Finally, as a proof of concept, they applied QLoRA fine-tuning on a single NVIDIA A100 GPU to three instruction-tuned models of similar size using the high-quality subset, converted each dialogue into progressive training samples predicting each teacher turn from preceding context, selected Mistral-7B via an auxiliary RoBERTa classifier trained on 2,134 safety-verified dialogues, and had 31 certified teachers rate its outputs against Claude Sonnet 4.5 and Gemini 2.0 Flash in a blinded, randomized setup analyzed with a linear mixed-effects model.
Why This Matters
Impact on research: ConvoLearn is positioned as the missing training-signal layer between inference-time prompting and evaluation-only benchmarks, and it is the comparison point in the paper's own table: MRBench (192 conversations, 8 dimensions) and SID (10,000 turns) are evaluation resources, while LearnLM and TeachLM are fine-tuning resources whose data and processes are not publicly available. ConvoLearn is presented as the only openly available resource combining dimension-level labels (6 dimensions, 21 subdimensions), conversation-level quality ratings, and public release. The paper also argues that LLM misalignment with pedagogy is structural — rooted in shared pretraining rather than model or prompt choice — which makes targeted, theory-driven training data necessary rather than merely useful.
Real-world applications:
- Building AI tutoring assistants for K-12 science that ask guiding questions instead of handing over answers.
- Training automated classroom observation tools, since a classifier trained on ConvoLearn correlates with expert-coded MQI and CLASS scores on authentic transcripts.
- Supporting teacher professional development by surfacing how often teachers use dialogic moves like scaffolding or metacognitive prompts.
- Enabling effectiveness-weighted or contrastive training workflows, because the full dataset deliberately retains dialogues across the entire quality spectrum rather than only the best ones.
Industry relevance: EdTech developers, LLM providers, and school districts evaluating AI tutors all need open, theory-grounded alignment data that proprietary systems do not expose. The result that a 7B open-weight model fine-tuned on 1,250 dialogues can approach a strong proprietary baseline suggests dialogic tutoring capability is attainable at modest scale, though Gemini 2.0 Flash's advantage indicates a larger gap remains at the frontier.
Future Directions
- Validate the remaining three dimensions: Construct validity is established for only cognitive engagement, formative assessment, and metacognition. Accountability, cultural responsiveness, and power dynamics depend on contextual, relational, and identity-based cues that transcript text cannot fully encode, and the authors suggest validation would require authentic student populations, longitudinal data, or multimodal signals.
- Move beyond imitation fine-tuning: The authors note that imitation training teaches surface form rather than pedagogical purpose, and propose preference-based feedback such as DPO or RLHF-style training using the full quality spectrum as a tractable extension.
- Broaden domain and population coverage: The dataset covers a single curricular domain and grade band, limiting generalizability across subjects, languages, and age groups, and it exhibits severe class imbalance with metacognition and cognitive engagement accounting for over half the data.
- Establish the link to learning outcomes: The connection between dialogic behavior and actual student learning remains unestablished in this work, and the authors argue deployment should evaluate impacts on learning and equity before scaling.
Target Audience
This paper is most useful to researchers and practitioners working at the intersection of learning sciences and AI: educational NLP researchers seeking theory-grounded training data, ML engineers building or aligning tutoring models, edtech teams designing dialogic AI assistants, and education researchers interested in how semi-synthetic data can be externally validated against authentic classroom observation instruments. Readers with a background in either instruction or model fine-tuning will find it accessible; readers without familiarity with correlation-based validation or parameter-efficient fine-tuning may want to skim the methodology sections. Educators and policymakers interested in the limits and risks of automated tutoring — including the authors' own ethical discussion of over-automation and uneven access — will also find the framing relevant.
Authors’ abstract
Despite their growing adoption in education, LLMs remain misaligned with the core principle of effective tutoring: the dialogic construction of knowledge. We introduce ConvoLearn, a dataset of 2,134 semi-synthetic tutor-student dialogues operationalizing six dimensions of dialogic tutoring grounded in knowledge-building theory, situated in a middle school Earth Science curriculum. We show that dimension-labeled dialogic training data captures meaningful pedagogical signal that generalizes beyond its semi-synthetic domain: scores from a classifier trained on ConvoLearn correlate significantly with expert-coded instructional quality in authentic classrooms across multiple subscales. As a proof of concept, we fine-tune Mistral-7B on ConvoLearn and show that dimension-level fine-tuning can steer a 7B open-weight model toward dialogic tutoring behavior that credentialed teachers rate as competitive with a strong proprietary baseline. With this work, we support the development of AI tutors capable of more dialogic interactions.