Research
Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback
Overview Research area: AI-assisted scholarly peer review, specifically the use of large language models (LLMs) to support human reviewers rather than replace them. Technical level: Beginner-Friendly.

- arXiv
- 2601.09182
- Published
- 2026-01-14
- Authors
- JungMin Yun, JuneHyoung Kwon, MiHyeon Kim, YoungBin Kim
AI summary
Overview
- Research area: AI-assisted scholarly peer review, specifically the use of large language models (LLMs) to support human reviewers rather than replace them.
- Technical level: Beginner-Friendly. This is a conceptual position paper. It contains no experiments, no datasets, and no quantitative evaluation of a built system; it proposes a design and an argument.
- Scope in one sentence: The paper argues that LLMs should be repositioned from automatic review generators to educational mentors and feedback critics for human reviewers, and it sketches a dual-system framework plus a five-principle rubric to support that shift.
What This Paper Is About
AI conferences are receiving far more submissions than their reviewer pools can handle well, which the authors call the "Reviewer Gap." This gap has two reinforcing parts: a Volume Gap (too many submissions for the available reviewing capacity) and a Quality Gap (not enough experienced domain experts, leading conferences to recruit junior or mandated reviewers without systematic training). The paper's goal is to argue that the fix is not faster automated review generation but mentoring and feedback that raise the competence of human reviewers over time.
Key Contributions
- A framing of the "Reviewer Gap" as two mutually reinforcing components — a Volume Gap (capacity constraints) and a Quality Gap (expertise constraints) — that together produce a vicious cycle of low-quality reviewing.
- A critique of existing LLM review approaches, arguing that direct LLM generation of initial reviews risks hallucinated information, inaccurate understanding of contributions and scholarly context, and superficial summarization, and that one-time LLM feedback interventions are insufficient without an educational framework.
- A unified rubric of five foundational principles for high-quality reviews — Fidelity, Clarity, Fairness, Proportionality, and Constructiveness — synthesized from requirements commonly used across major conferences (AAAI, NeurIPS, and others cited as "Review 2025").
- A dual-system framework: (i) an LLM-assisted reviewer mentoring system with a three-stage curriculum (Guided Recognition, Review Refinement Practice, Full Simulation) ending in a voluntary Reviewer Certification, and (ii) an LLM-assisted reviewer feedback system that runs after a draft is submitted (Detection and Cross-Verification, Evidence-based Feedback Generation, then a Reliability Testing check before any feedback is shared).
Main Findings
- Growth is outpacing review capacity: The paper reports that conferences such as ICLR and ACL show annual submission increases of 20–30%, and that NeurIPS received more than 27,000 submissions in 2025, roughly a tenfold increase over the past decade.
- Two reinforcing gaps: The Volume Gap produces excessive assignments per reviewer and review fatigue, leading to superficial commentary and perfunctory assessments; the Quality Gap arises because expert growth lags submission growth, and expanded reviewer pools (for example via ARR and NeurIPS author-participation mandates) widen expertise disparities.
- Automated generation is the wrong target: Because reviewing demands expert judgment — assessing contributions, contextualizing within the literature, and evaluating methodological validity — the authors argue current LLMs cannot fully substitute for that expertise, and that inexperienced reviewers relying on LLM drafts without critical examination may reproduce misunderstandings and biases.
- Feedback alone is not enough: A feedback-oriented approach implemented as a one-time intervention risks reviewers accepting incomplete or biased suggestions uncritically, reinforcing low-quality practices; it must evolve into a long-term training mechanism.
- The rubric is the anchor: The five principles (Fidelity, Clarity, Fairness, Proportionality, Constructiveness) are proposed as the shared reference points that let an LLM consistently evaluate and support review improvement, aiming to cultivate better reviewers rather than merely better reviews.
- Design safeguards against the main objections: The authors address concerns about deskilling, automation bias, bias amplification, and review homogenization by pointing to the system's optional refinement structure (revision is never mandatory, and reviewers keep full autonomy) and its explicit educational purpose. They also state that the LLM checks the form in which subjective opinions are expressed, not the substantive validity of reviewers' judgments.
- Stated limitation: The framework focuses on review text quality and does not yet incorporate analysis of supplementary materials such as code repositories or datasets.
Methodology in Plain English
Because this is a position paper, the "method" is argumentation and system design rather than experimentation. The authors first diagnose the problem by surveying reported submission trends and citing existing literature on LLM review systems and reviewer training. They then define what a good review should contain by consolidating criteria used at major conferences into five principles. Using those principles as a rubric, they propose two complementary systems: a training pipeline that walks reviewers through recognizing good and bad reviews, revising deliberately flawed drafts, and finally writing a full review under LLM supervision; and a post-draft pipeline where an LLM acts as a critic, cross-checks the review against the paper, generates evidence-backed feedback, screens that feedback for tone and usefulness, and privately shares it with reviewers and Area Chairs. They also propose that Area Chairs submit meta-feedback on the LLM's suggestions so that corrected cases accumulate as training data for continual refinement of the system.
Why This Matters
Peer review is the credibility mechanism of academic publishing, and the paper argues that the current strain is structural rather than individual: it affects conference credibility and researcher motivation, not just reviewer workloads. The paper's distinctive move is to treat review quality as an educational problem — a matter of cultivating competence — rather than a generation-speed problem, and to reserve final judgment for humans.
Real-world applications suggested by the framework:
- Conference reviewer onboarding: A voluntary, "safe-to-fail" simulated training environment where new reviewers practice on example reviews and flawed drafts before handling real submissions.
- Pre-submission review triage: An automated check that flags vague or unsupported criticisms in a submitted review draft and returns specific, evidence-based suggestions for the reviewer.
- Area Chair oversight: Private sharing of system-generated feedback and reliability-tested suggestions with Area Chairs, who can also supply meta-feedback to correct LLM errors or biases.
- Certification signaling: A voluntary Reviewer Certification acting as a positive signal of a reviewer's investment in community quality, without functioning as a mandatory gate.
Industry relevance: the proposal targets the operational review pipelines of large AI conferences and, more broadly, any venue or journal facing submission growth that outpaces expert supply. It also implies a design philosophy for deploying LLMs in expert workflows — as augmenting assistants with human authority retained — that transfers to other high-stakes evaluation settings.
Future Directions
- Extending coverage to the full submission package, including code repositories and datasets, and incorporating reproducibility verification rather than only review-text quality.
- Building and evaluating the actual systems. The paper presents no prototype results, no user study, and no measured effect on review quality; whether the mentoring curriculum and feedback critic work as claimed remains untested.
- Validating the rubric empirically, including whether the five principles can be applied consistently by an LLM across papers and fields.
- Studying the risks the authors acknowledge, such as automation bias, reviewer deskilling, and homogenization of reviews, and whether the optional-refinement structure genuinely mitigates them.
- Operationalizing the collaborative improvement loop, in which Area Chair meta-feedback is accumulated as training data to continuously refine the system in alignment with conference standards.
Target Audience
Conference organizers, program chairs, and Area Chairs who manage reviewer recruitment, training, and quality control; peer-review researchers and meta-science scholars interested in LLM integration in scholarly evaluation; reviewers in training who want a structured account of what constitutes a high-quality review; and AI researchers and product teams building LLM tooling for expert-in-the-loop evaluation workflows. Readers seeking empirical results on LLM-assisted review performance will not find them here, as the paper is a position piece.
Authors’ abstract
The rapid expansion of AI research has intensified the Reviewer Gap, threatening the peer-review sustainability and perpetuating a cycle of low-quality evaluations. This position paper critiques existing LLM approaches that automatically generate reviews and argues for a paradigm shift that positions LLMs as tools for assisting and educating human reviewers. We define the core principles of high-quality peer review and propose two complementary systems grounded in these foundations: (i) an LLM-assisted mentoring system that cultivates reviewers' long-term competencies, and (ii) an LLM-assisted feedback system that helps reviewers refine the quality of their reviews. This human-centered approach aims to strengthen reviewer expertise and contribute to building a more sustainable scholarly ecosystem.