Research
Compact Prompting in Instruction-tuned LLMs for Joint Argumentative Component Detection
Overview Research area: Natural Language Processing, specifically Argument(ation) Mining (AM) and its core subtask, Argumentative Component Detection (ACD). Technical level: Intermediate — readers wil

- arXiv
- 2603.03095
- Published
- 2026-03-03
- Authors
- Sofiane Elguendouze, Erwan Hain, Elena Cabrio, Serena Villata
AI summary
Overview
Research area: Natural Language Processing, specifically Argument(ation) Mining (AM) and its core subtask, Argumentative Component Detection (ACD).
Technical level: Intermediate — readers will benefit from familiarity with instruction-tuned LLMs, fine-tuning, sequence labeling (BIO tagging), and standard evaluation metrics such as macro-F1.
Scope: The paper reframes ACD as a text-generation task and evaluates compact instruction-based prompting with fine-tuned open-weight LLMs on three argument-annotated benchmarks against CRF, heuristic, multi-task, and encoder-based baselines.
What This Paper Is About
Argumentative Component Detection requires finding where argumentative spans begin and end in raw text (segmentation) and deciding whether each span is a claim or a premise (classification). Most prior work simplifies this by assuming argumentative units are already delimited, turning ACD into plain classification, or by using multi-stage pipelines that separate segmentation from labeling. The authors instead ask an instruction-tuned LLM to reproduce the input text with XML tags inserted around argumentative components, so that segmentation and classification are produced jointly in a single generation step.
Key Contributions
-
A generative reformulation of ACD. The task is cast as text-to-structure generation: the model receives plain, unsegmented text and outputs the same text with
<claim>...</claim>and<premise>...</premise>tags, so boundary detection and component labeling happen in one pass rather than in a pipeline. -
A compact instruction-prompting recipe and data conversion. Corpora originally annotated with the BIO scheme are converted into XML-tagged text, and each instance becomes an input–output pair of plain text and its tagged counterpart, used for instruction tuning. A single reference output may contain multiple argumentative components.
-
Systematic fine-tuning of open-weight LLMs. The authors fine-tune GPT-2-XL-1.5B, OPT-1.3B, OPT-6.7B, Mistral-7B-v0.3, and Llama-3-8B-Instruct, and additionally train encoder-based RoBERTa and DeBERTa baselines under traditional token-level BIO classification, evaluating both on the Persuasive Essays benchmark and on a merged corpus of all three datasets.
-
Best-in-table results plus a qualitative error analysis. The best model reaches a macro-F1 of 0.8778 on Persuasive Essays, above every reported baseline and within 8.2 × 10⁻³ of the reported human upper bound. The authors also document four recurring qualitative phenomena, including cases where the model's label arguably improves on the gold annotation. Code and datasets are stated to be openly available in a GitHub repository.
Main Findings
-
Strongest overall result: Llama-3-8B fine-tuned on Persuasive Essays achieved the highest macro-F1 (0.8778) and accuracy (0.9004) among all methods, outperforming the strongest feature-engineered CRF (0.8670) and the reported human upper bound (0.8860) gap of only 8.2 × 10⁻³.
-
Other Persuasive Essays results: OPT-6.7B reached 0.8518 macro-F1 (accuracy 0.8856) and GPT-2-1.5B reached 0.8521 (accuracy 0.8804) — both below Llama-3-8B, but each above the multi-task MT-all baseline from Morio et al. (0.7566).
-
Baselines on Persuasive Essays: the heuristic baseline scored 0.642, CRF with features 0.8670, the CRF from Goudas et al. (2014) 0.4237 (included only for historical comparison, since it was run on a different dataset — Greek social media text), DeBERTa-v3 0.7112, and RoBERTa 0.6933.
-
Merged-corpus setting is harder: training and testing on the combination of all three datasets caused a performance drop of 0.0956 relative to the Persuasive Essays setting. The best merged-model result was OPT-6.7B at 0.7822 macro-F1 (accuracy 0.8076); other merged results were GPT-2-1.5B 0.7684 (accuracy 0.7955), OPT-1.3B 0.7679 (accuracy 0.7971), Llama-3-8B 0.7667 (accuracy 0.8132), and Mistral-7B 0.7652 (accuracy 0.8016).
-
Encoder models degrade sharply across domains: RoBERTa and DeBERTa-v3 fell to 0.48 and 0.49 macro-F1 on the merged setting, against 0.6933 and 0.7112 on Persuasive Essays. The authors interpret this as traditional encoders struggling more with cross-domain variability than generative models.
-
Class-wise behavior (BIO analysis): on Persuasive Essays, Llama-3-8B scored B-C 0.84, I-C 0.80, B-P 0.90, I-P 0.88, O 0.98 (macro-F1 0.88); OPT-6.7B and GPT-2-1.5B each scored 0.78 B-C, 0.76 I-C, 0.87 B-P, with macro-F1 0.85. Premise boundaries were consistently detected more accurately than claim boundaries across generative models, while the O class reached up to 0.98.
-
Merged class-wise behavior is more uniform: OPT-6.7B scored B-C 0.77, I-C 0.73, B-P 0.74, I-P 0.74, O 0.93 (macro-F1 0.78), and the other merged models clustered near 0.77–0.78, suggesting a more generalized representation rather than a claim/premise asymmetry.
-
Encoder class-wise weaknesses: DeBERTa-v3 and RoBERTa on Persuasive Essays scored much lower on boundary tags, e.g., RoBERTa B-C 0.59 and I-C 0.68, DeBERTa-v3 B-C 0.6 and B-P 0.58, indicating difficulty with precise span delimitation.
-
Four qualitative phenomena were observed: (1) argument type refinement, where the model assigns a different but arguably more appropriate label, exposing possible annotation inconsistencies; (2) surface-level textual changes, such as small grammatical or lexical edits that break exact string matching; (3) discovery of unannotated argumentative components, counted as false positives under strict evaluation; and (4) hallucination, where words are added or altered despite instructions to reproduce the input verbatim, which misaligns spans and lowers scores.
-
Hallucination mitigation was incomplete: lowering temperature, restricting nucleus sampling, strengthening prompt constraints, and emphasizing verbatim reproduction did not fully eliminate the problem, which the authors suggest is intrinsic to autoregressive models optimized for fluent continuation rather than strict copying.
Methodology in Plain English
The authors treat detection as a rewriting exercise. For each dataset annotated with BIO labels, they convert the annotations into XML-style tags placed inside the original text, producing training pairs where the input is the raw sentence or passage and the target is the same text with <claim> and <premise> tags inserted. Models are then fine-tuned on these pairs so that, at inference time, given plain unsegmented text, they emit the tagged version — generating boundaries and labels together.
Three corpora are used: USElecDeb60To16 (annotated 1960–2016 U.S. presidential debate transcripts, described as the largest argument-annotated dataset to date, with 29k claims and 26k premises), Persuasive Essays (402 English essays from essayforum.com, with 2257 claims and 3832 premises), and Web Discourse (user-generated web texts on six controversial education-related topics, with 195 claims and 538 premises). A merged corpus combining all three contains 31.5k claims and 30.3k premises. The datasets differ in size, style, structural clarity, and discourse complexity, giving a heterogeneous evaluation setting.
Decoding is configured for near-determinism: temperature 0.01 and nucleus sampling with top-p = 0.1. Longer inputs are split into chunks of up to 1024 tokens. Every model is fine-tuned for at least 10 epochs, and the checkpoint with the highest macro-F1 on the validation set is selected; all experiments use an 80/10/10 train–validation–test split. The encoder baselines (RoBERTa, DeBERTa) are instead trained for token-level classification with the BIO scheme, for at least 10 epochs, learning rate 1e-4, and a maximum input length of 64 tokens.
The authors deliberately restricted themselves to open-weight models — despite proprietary systems such as GPT-4o being available — for reproducibility, transparency, controlled fine-tuning, and architectural inspection, as well as cost and data-privacy constraints in real-world deployment. Results are reported for two configurations: the merged corpus, and Persuasive Essays alone for comparison with established baselines.
Why This Matters
The paper provides evidence that joint segmentation-and-labeling of arguments can be learned as a single generation task instead of a pipeline, and that fine-tuned open-weight LLMs can exceed strong feature-engineered systems such as CRFs on this benchmark. It also raises a methodological point: strict span-level evaluation may penalize predictions that are semantically equivalent or arguably better than the gold annotation, which matters for how ACD progress is measured.
Real-world applications named or implied in the paper:
- Opinion analysis in large-scale deliberation, where arguments in public discussion need to be identified and decomposed automatically.
- Decision support, drawing on structured argumentative components extracted from source documents.
- Educational technologies, for instance analyzing argumentative essays of the kind contained in the Persuasive Essays corpus.
- Argument analysis in specialized domains such as legal decisions, following prior AM work on European Court of Human Rights decisions cited in the paper's related work.
For industry, the findings suggest that a single generative model can replace multi-stage ACD pipelines, which reduces pipeline complexity and the cumulative segmentation-then-classification error the authors describe. The exclusive use of open-weight models is also positioned as practical for settings with cost or data-privacy constraints. The paper cautions that argumentation technologies may be misused in contexts such as political persuasion or large-scale rhetorical analysis, particularly when combined with generative capabilities, and that hallucinations or small alterations to input text could distort meaning in sensitive contexts.
Future Directions
- Constrained decoding and hybrid extractive–generative approaches to enforce verbatim input fidelity, since temperature reduction, top-p restriction, prompt tightening, and explicit verbatim instructions did not eliminate hallucination.
- Evaluation protocols that account for semantic equivalence, so that minor surface edits and arguably better labels are not automatically scored as errors.
- Extending the generative formulation beyond claim and premise detection to relation identification, stance modeling, and full argument graph construction, which the authors explicitly list as unexplored.
- Investigating whether model predictions can surface annotation noise or ambiguities in existing corpora, given the observed argument type refinements and unannotated components.
Target Audience
NLP and computational-argumentation researchers, particularly those working on argument mining, sequence labeling, and structured generation; practitioners who need to extract claims and premises from raw, unsegmented text; and evaluators or dataset builders interested in annotation quality and the limitations of strict span-level scoring with generative models.
Authors’ abstract
Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems.