Research
Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
Overview Research area: Natural Language Processing — Automated Essay Scoring (AES), with a focus on learning under limited labeled data (few-shot / semi-supervised settings). Technical level: Advance

- arXiv
- 2602.01747
- Published
- 2026-02-02
- Authors
- Hongseok Choi, Serynn Kim, Wencke Liermann, Jin Seong, Jin-Xia Huang
AI summary
Overview
- Research area: Natural Language Processing — Automated Essay Scoring (AES), with a focus on learning under limited labeled data (few-shot / semi-supervised settings).
- Technical level: Advanced. The paper assumes familiarity with transformer fine-tuning, low-rank adaptation, model calibration/score-distribution alignment, and semi-supervised self-training.
- Scope (one sentence): The paper proposes three techniques — Two-Stage fine-tuning with low-rank adaptations, Score Alignment, and uncertainty-aware self-training — and integrates them into the DualBERT AES model, evaluating them on ASAP++ with additional generalization checks on TOEFL11 and ELLIPSE.
What This Paper Is About
Automated Essay Scoring can make writing assessment faster and more scalable, but building a reliable scoring model normally requires large amounts of human-scored essays, which are expensive and hard to obtain. The paper's goal is to make AES work well when labeled data is scarce, and to keep it strong even when plenty of labeled data is available. To do this, the authors combine three complementary techniques and apply all of them on top of an existing model, DualBERT.
Key Contributions
- Two-Stage fine-tuning strategy using low-rank adaptations (LoRA). The approach is designed to better adapt an AES model to the specific target prompt's essays.
- Score Alignment technique. A method intended to improve consistency between the distribution of predicted scores and the distribution of true scores.
- Uncertainty-aware self-training on unlabeled data. Unlabeled essays are pseudo-labeled and added to the training set, with the uncertainty-aware design intended to limit the propagation of label noise from incorrect pseudo-labels.
- Integration and generalization study. All three techniques are implemented on DualBERT and evaluated primarily on ASAP++ and additionally on TOEFL11 and ELLIPSE to test whether the techniques generalize beyond one dataset.
Main Findings
- All three techniques help in the low-data regime: In the 32-data setting on ASAP++, the abstract reports that each of the three key techniques improves performance.
- Combined techniques approach full-data performance: Their integration reaches 91.2% of the full-data performance, where the full-data model was trained on approximately 1,000 labeled samples.
- Score Alignment is consistently beneficial: The abstract states that Score Alignment improves performance in both limited-data and full-data settings, rather than only in the scarce-data case.
- State-of-the-art claim: Integrated into DualBERT, Score Alignment achieves state-of-the-art results in the full-data setting on ASAP++.
- Generalization is examined but not quantified in the abstract: The authors also evaluate the techniques on TOEFL11 and ELLIPSE, but the abstract does not report the resulting numbers, so the size or direction of the cross-dataset effects cannot be stated from the abstract alone.
Methodology in Plain English
The researchers start from DualBERT, an existing essay-scoring model, and add three layers of improvement.
First, instead of fine-tuning the whole model in one go, they use a two-stage procedure that relies on low-rank adaptations — small, efficient add-on parameters that let the model adjust to the particular essay prompt without retraining everything. This makes adaptation to a new prompt cheaper and, per the paper, more effective.
Second, they add a Score Alignment step. Scoring models can predict the right average score while still being wrong about how scores are spread across the scale; this technique pushes the predicted score distribution to line up with the true one.
Third, they use unlabeled essays through self-training: the model generates its own labels for unlabeled essays, and those pseudo-labeled examples are added to the training set. Because wrong pseudo-labels would teach the model bad habits, the method is uncertainty-aware — it takes the model's confidence into account so that noisy pseudo-labels are less likely to contaminate training.
They test this combined system most extensively on ASAP++, including a deliberately tiny 32-example setting and a full-data setting of roughly 1,000 labeled essays, and they run the same techniques on TOEFL11 and ELLIPSE to see whether the benefits transfer.
Why This Matters
- Research impact: The paper targets a well-known bottleneck in AES — labeled data scarcity — and shows that a combination of parameter-efficient fine-tuning, output-distribution alignment, and noise-aware semi-supervised learning can be layered onto an existing model. If the Score Alignment result holds up, it is a comparatively simple addition that pays off in both data-poor and data-rich conditions, which is unusual and worth attention from researchers working on scoring and calibrated regression generally.
- Real-world applications:
- Large-scale writing assessment in schools or standardized testing, where per-prompt human scoring is costly.
- Adaptive learning platforms that need to score student essays on newly introduced prompts with few or no human-graded examples.
- Second-language writing evaluation, relevant given the additional evaluation on TOEFL11 (a learner-English corpus).
- Rapid deployment of new essay prompts, where a model must adapt to a prompt that has little or no labeled data yet.
- Industry relevance: Any organization that grades free-text responses at volume — testing companies, edtech vendors, MOOC providers, and internal training/assessment teams — has an incentive to reach acceptable accuracy with far fewer human-scored examples, since annotation is the dominant cost. The 91.2%-of-full-data result at roughly 32 labeled examples, if it generalizes, directly reduces that cost. The abstract does not discuss deployment constraints, fairness across demographic groups, or cost of inference, so those remain open for practitioners.
Future Directions
- Report and analyze the TOEFL11 and ELLIPSE results in detail, since the abstract only states that these datasets were used to examine generalizability, not what the outcomes were.
- Isolate the contributions more finely, including how each technique interacts with the others and whether the two-stage LoRA fine-tuning provides the same benefit for models other than DualBERT.
- Test whether the approach transfers to other AES models and scoring schemes, since the paper implements the techniques specifically on DualBERT and reports state-of-the-art only on ASAP++.
- Investigate robustness and fairness of pseudo-labeling at very small sample sizes, given that self-training with noisy labels is the component most likely to introduce systematic bias, and the abstract does not address this.
Target Audience
Researchers and practitioners working on automated essay scoring, educational NLP, and semi-supervised or low-resource text regression. It is also relevant to machine learning engineers interested in combining parameter-efficient fine-tuning with distribution alignment and uncertainty-aware self-training, and to edtech or assessment teams evaluating whether AES can be deployed without large annotated corpora. Readers without background in transformer fine-tuning and semi-supervised learning may find the techniques difficult to follow from the abstract alone.
Authors’ abstract
Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world settings, the extreme scarcity of labeled data severely limits the development and practical adoption of robust AES systems. This study proposes a novel approach to enhance AES performance in both limited-data and full-data settings by introducing three key techniques. First, we introduce a Two-Stage fine-tuning strategy that leverages low-rank adaptations to better adapt an AES model to target prompt essays. Second, we introduce a Score Alignment technique to improve consistency between predicted and true score distributions. Third, we employ uncertainty-aware self-training using unlabeled data, effectively expanding the training set with pseudo-labeled samples while mitigating label noise propagation. We implement the above three key techniques on DualBERT. We conduct extensive experiments on the ASAP++ dataset, and additionally evaluate the proposed techniques on two other datasets, TOEFL11 and ELLIPSE, to examine their generalizability. In the 32-data setting on ASAP++, all three key techniques improve performance, and their integration achieves 91.2% of the full-data performance trained on approximately 1,000 labeled samples. In addition, the proposed Score Alignment technique consistently improves performance in both limited-data and full-data settings: e.g., it achieves state-of-the-art results in the full-data setting on ASAP++ when integrated into DualBERT.