Research
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Overview Research area: Natural Language Processing, specifically the evaluation and design of large language model (LLM) systems that assist or automate scientific peer review. Technical level: Inter

- arXiv
- 2609.39027
- Published
- 2026-09-30
- Authors
- Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen, Han Chen, Tianyi Zhou, Dawei Zhou
AI summary
Overview
Research area: Natural Language Processing, specifically the evaluation and design of large language model (LLM) systems that assist or automate scientific peer review.
Technical level: Intermediate. The paper assumes familiarity with LLM benchmarking concepts (prompting protocols, score prediction, correlation metrics), but its framing of "rhetorical robustness" is explained from first principles.
Scope: The paper argues that trustworthy AI reviewers must be both stable when the same science is reworded and still able to distinguish different papers, introduces a 1,260-manuscript benchmark called RobustReview to measure this, and proposes a dual-branch reviewer, SciCore, that averages a full-manuscript judgment with a judgment of an extracted "science core."
What This Paper Is About
LLMs are increasingly used to generate manuscript feedback and support accept/revise decisions, and prior evaluations mostly ask whether AI reviews look useful or agree with human scores. The problem the authors identify is that a reviewer can assign different judgments to manuscripts that report the same science in different wording, which would reward rhetorical optimization over actual scientific improvement. The goal is to formalize this blind spot as "rhetorical robustness," build a controlled benchmark to measure it, and design a reviewer architecture that improves it.
Key Contributions
-
A formalization of Rhetorical Robustness as a joint requirement: within-paper stability across content-preserving rewrites plus between-paper discrimination, with human alignment treated as a separate evaluation dimension rather than a substitute.
-
RobustReview, a controlled full-manuscript benchmark built from 60 anonymized ICLR 2026 submissions, expanded with 10 rhetorical conditions produced by two independent LLM systems into 1,260 full manuscripts total, and used to evaluate 30 reviewer configurations spanning general-purpose LLMs, specialized scientific-review models, and agentic review systems.
-
The identification of "false robustness" as a central measurement failure, where low rewrite-induced score drift coincides with score collapse across papers, meaning a reviewer appears stable only because it barely differentiates papers at all.
-
SciCore, a dual-branch framework that averages a conventional full-manuscript judgment with a content-normalized judgment made on an extracted, structured science core (reported problem, claims, methods, assumptions, evidence, results, contributions, reproducibility information, and limitations).
Main Findings
-
Content-focused prompting alone is not reliable. The Persistent protocol, which repeatedly instructs the reviewer to base scientific judgments on substantive content rather than rhetoric, does not consistently improve robustness across backbones. Only GPT-5.5 improves on all five robustness metrics relative to Standard. For Claude Sonnet 5, Persistent reduces MAD and Drift SD but lowers ICC, SPR, and discriminability.
-
False robustness is widespread. Gemini-3.5-Flash-Lite under Standard achieves the lowest MAD among the existing configurations, at 0.205, but its ICC is only 0.199 and its discriminability is 0.511, close to chance. DeepReviewer and OpenJudge also show relatively low drift yet attain ICC values of only 0.161 and 0.239 and discriminability values of 0.537 and 0.526.
-
Stability and human alignment rank reviewers differently. Within GPT-5.5, Strict gives the strongest human-alignment profile among existing configurations (Human MAE 1.078, Spearman 0.529), whereas Persistent has weaker human alignment but higher ICC, SPR, and discriminability. GPT-5.5 under Persistent has a higher MAD of 0.476 but the strongest ICC, SPR, and discriminability among the 30 configurations.
-
SciCore leads on joint stability-discrimination while staying competitive on alignment. In the main comparison, SciCore reaches ICC 0.775, SPR 0.652, and discriminability 0.726, the lowest Human MAE at 1.072, and the second-highest human Spearman correlation at 0.488. It does not achieve the lowest MAD or Drift SD (0.476 and 0.687 respectively), and its human Spearman correlation remains below GPT-5.5 under Strict.
-
Direct review of the science core beats reconstruction in this pipeline. Relative to Manuscript-Standard (MAD 0.598, Drift SD 0.956), ReconstructReview worsens MAD to 0.651 and Drift SD to 1.101, while Core-Standard improves all seven metrics over ReconstructReview.
-
Extraction alone does not solve the review-policy problem. Core-Persistent worsens all five robustness metrics relative to Core-Standard, whereas Core-Adapted achieves the lowest MAD (0.292) and Drift SD (0.650) and the highest ICC (0.751) and SPR (0.710) among the core-only configurations.
-
The two branches are complementary. Fusing Core-Adapted with Manuscript-Strict reduces MAD from 0.766 to 0.476 and increases ICC from 0.632 to 0.775, improving all five robustness metrics over Manuscript-Strict, and also improves human alignment, ICC, and discriminability over Core-Adapted.
-
Extracted science cores are stable yet paper-specific. For GPT-5.5, matched cores have a mean cosine similarity of 0.978, while cross-paper similarity ranges from 0.616 to 0.619. Each matched cell contains 600 pairs and each control cell 35,400 comparisons per rewrite producer.
-
Some criteria expose false robustness directly. Core-Standard and Core-Strict assign a constant confidence score (MAD 0.000, SPR 0.000), and Core-Persistent produces almost no variation (MAD 0.003), whereas Core-Adapted restores paper-level variation and raises confidence SPR to 0.394.
-
Cross-backbone transfer is partial. Core-Adapted improves ICC for all three tested backbones. GPT-5.5 improves on all five robustness metrics, GPT-5-mini improves MAD, Drift SD, and ICC but weakens SPR and discriminability, and GLM-5.2 improves all seven reported metrics. Human alignment does not improve consistently.
-
Equal weighting is a reasonable default. With the prespecified α = 0.5 science-core weight under Manuscript-Strict, larger science-core weights tend to improve robustness while weakening human alignment overall; the prespecified configuration is reported as Pareto non-dominated among the evaluated protocol-weight combinations on the seven reported point estimates.
Methodology in Plain English
The authors first define what a trustworthy reviewer should do: give roughly the same scientific judgment when a paper is rewritten in a way that preserves its reported science, while still giving different papers different scores. They then build a test set to measure this. Starting from 60 anonymized ICLR 2026 submissions that had matched arXiv LaTeX sources, they stratify papers by mean human overall-assessment rating and randomly sample 10 papers from each of six score intervals. Each paper is kept in its original form and also rewritten under 10 rhetorical conditions — six that alter a single dimension (novelty stance, scope framing, evidence framing, contribution salience, technical register, or linguistic complexity) and four complex ones (a joint multi-dimension rewrite, two or three recursive rewrite rounds, and a reviewer-guided rewrite based on model feedback). Each condition is independently produced by GPT-5.5 and Claude Opus 4.8, giving two variants per paper per condition: 60 originals plus 1,200 variants, or 1,260 manuscripts.
They then run 30 reviewer configurations over this corpus, holding the scoring criteria fixed within each configuration. General-purpose models are prompted under three protocols (Standard, used as the ICLR-style baseline; Strict, which raises the evidentiary threshold; and Persistent, which repeatedly tells the reviewer to focus on content over presentation). Specialized review models and agentic systems are run with their native procedures. Every configuration is scored on seven metrics: two direct within-paper stability measures (MAD and Drift SD), three joint stability-discrimination measures (ICC, SPR, and discriminability), and two human-alignment measures (Human MAE and Spearman correlation).
Finally, they design SciCore. One branch reviews the full manuscript under the Strict protocol. The other uses an LLM to extract a structured science core — central idea and claims, problem formulation, mathematical formulations and derivations, methods and assumptions, evidence, reported results, author-stated contributions, reproducibility information, and stated limitations — written in objective third-person language with tables transcribed and uncertainty recorded rather than inferred. That record, not the manuscript, is then reviewed with an adapted protocol. The final score is the unweighted arithmetic mean of the two branch scores.
Why This Matters
The paper reframes what "good" AI reviewing means. Prior evaluations largely measure agreement with human scores or review usefulness, but the authors show that human alignment and rhetorical robustness favor different reviewer configurations, so agreement with humans is not evidence that a reviewer is resistant to wording changes. A reviewer that can be nudged by rewriting alone would reward rhetorical optimization over scientific improvement, which matters for which work is accepted, revised, and disseminated — a point the authors make about the broader scientific record.
Real-world applications:
- Conference and journal peer review. The benchmark uses real ICLR 2026 submissions and asks whether automated judgments survive visible, meaning-preserving edits to titles, abstracts, and full manuscripts, which is directly relevant to submission screening and review triage.
- Reviewer-assistance tools. The science-core branch shows a way to produce a content-normalized assessment alongside a conventional one, giving authors and editors a second view that is less coupled to presentation.
- Manuscript preparation and revision. Understanding which rhetorical dimensions move automated judgments (novelty stance, evidence framing, technical register, and others) helps authors and editors separate legitimate clarity improvements from presentation changes that should not alter scientific merit.
- Auditing deployed AI judges. The "false robustness" diagnosis — low drift combined with near-chance discrimination — gives a concrete failure signature to check for when evaluating any scoring system, not just reviewers.
Industry relevance: Any organization deploying LLM-based judges for scientific, technical, or grant evaluation faces the same failure mode. The paper's ablation showing that content-focused prompting does not consistently help, and that a constant-output branch can look perfectly stable, is a warning for teams that validate their scorers only on stability.
Future Directions
- Extending the dual-branch fusion across backbones. The cross-backbone experiments only test the science-core branch against Manuscript-Standard; the paper explicitly states that these experiments do not test the final dual-branch fusion across backbones, so branch-level transfer is only partially established.
- Improving human alignment without sacrificing robustness. SciCore's human Spearman correlation (0.488) remains below GPT-5.5 under Strict (0.529), and larger science-core weights tend to weaken human alignment, so a fusion that gains robustness without an alignment loss remains open.
- Tuning versus fixing the fusion weight. The paper reports a prespecified equal-weight default and a sensitivity sweep over α, but a principled way to choose weights per protocol or per criterion is not settled.
- Diagnosing and fixing output collapse. Since Core-Standard and Core-Strict assign constant confidence scores and Core-Persistent produces almost none, methods for detecting and preventing near-constant outputs on secondary criteria are an unresolved problem.
Target Audience
This paper is most useful for researchers building and evaluating LLM-based judges and review systems, for program chairs and editors considering automated review support, and for benchmarking practitioners interested in robustness evaluation. It also suits NLP researchers studying prompt sensitivity and invariance, since the science-core branch is framed as an invariant-representation argument with an explicitly empirical validation requirement. Readers seeking a fully worked-out deployed system will find the work framed as a benchmark plus a design proposal, with several components (cross-backbone fusion, weight selection) left as open questions.
Authors’ abstract
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.