Research
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
Overview Research area: Natural Language Processing / in-context learning (ICL) in large language models; specifically the interaction between pre-trained label semantics and few-shot demonstration le

- arXiv
- 2511.21038
- Published
- 2025-11-26
- Authors
- Anantha Padmanaban Krishna Kumar
AI summary
Overview
- Research area: Natural Language Processing / in-context learning (ICL) in large language models; specifically the interaction between pre-trained label semantics and few-shot demonstration learning.
- Technical level: Intermediate. The paper uses standard classification benchmarks and simple prompt-based evaluation, but its central argument depends on a custom metric decomposition (truth, prior, and prompt alignment) that readers should be comfortable interpreting.
- Scope (one sentence): The paper tests whether eight open-source LLMs between 1B and 12B parameters can learn inverted label mappings from few-shot demonstrations, or whether their pre-trained label semantics are immovable anchors.
What This Paper Is About
In-context learning lets a frozen language model adapt to a task from a handful of labeled examples, but it is unclear whether the model is genuinely learning a new input-to-label mapping or simply sharpening a classifier it already has from pre-training. This paper tests the two theories head-on by giving models demonstrations whose labels are systematically flipped (for example, calling positive reviews NEG), and measuring whether the model can become an accurate "anti-semantic" classifier. The author argues that across eight classification tasks and eight models in the 1–12B range, ICL only refines existing semantic priors and never overrides them.
Key Contributions
- A three-way decomposition of ICL behavior. The paper separates ICL predictions into three alignment metrics: truth alignment (accuracy), prior alignment (agreement with the zero-shot prediction), and prompt alignment (agreement with the demonstrated mapping).
- The semantic override rate. A new metric defined as the probability that a prediction is both correct and consistent with the inverted mapping, under inverted demonstrations. This captures genuine semantic flexibility rather than label-following alone.
- A systematic natural-vs-inverted comparison across eight models and eight tasks. The study spans the LLaMA, Mistral, Qwen, and Gemma families, from 1B to 12B parameters, including a base-versus-instruction-tuned comparison (LLaMA 3.1 8B), at demonstration counts k ∈ {1, 2, 4, 8}.
- An empirical claim of exactly zero override. The semantic override rate is reported as exactly 0.0% across all 320 experimental conditions (8 models × 8 tasks × 5 seeds), not merely near zero.
Main Findings
- Natural ICL improves accuracy. For LLaMA-3.1-8B-Instruct, average accuracy rises from 69.5% zero-shot to 79.3% at k = 8. Gains are largest where zero-shot priors are weakest: QQP jumps 37.8 points (40.6% → 78.4%), while NLI tasks gain 14–17 points, and sentiment classification adds only 2–3 points despite starting above 90%.
- Inverted ICL degrades accuracy. Average accuracy falls to 52.8% at k = 8. Sentiment classification loses over 40 points and topic classification drops 32 points. The exception is QQP, where inverted ICL reaches 71.6%—above its weak zero-shot baseline (40.6%) but still 7 points below natural ICL.
- ICL works by prior refinement, not replacement. Under natural demonstrations, predictions stay tightly coupled to zero-shot behavior: SST-2 achieves 92.5% accuracy at k = 8 while keeping 89.8% alignment with zero-shot predictions, with 86.4% of examples both correct and prior-consistent. Even on QQP, where accuracy rises by 37 points, the model still agrees with its zero-shot prior on 34.9% of examples.
- The semantic override paradox. Models partially follow inverted demonstrations—prompt agreement reaches 51.6% on IMDB and 52.6% on SST-2—but the semantic override rate stays exactly zero. No example is simultaneously correct and consistent with the inverted mapping.
- Degradation is monotonic with more examples. Under inversion, tasks that barely degrade at k = 1 (losing 5–10 points on sentiment) collapse at k = 8 (losing more than 40 points); each added inverted example strengthens the conflict between demonstrated and pre-trained semantics.
- The pattern holds across scale. The behavior is consistent across the twelvefold parameter range (1B to 12B), with no evidence that larger models in this range overcome the anchors. The base-vs-instruct LLaMA 3.1 8B comparison is included to isolate the effect of fine-tuning on semantic rigidity.
- QQP is anomalously weak, yet still anchored. QQP shows unusually weak zero-shot performance across all models (24.7%–85.0%), and models still cannot override its label semantics despite the weak prior.
Methodology in Plain English
The author treats each model as a prompt-induced classifier and runs three prompting conditions on every test item: zero-shot (instructions only), natural k-shot (correct labels), and inverted k-shot (labels systematically flipped by a permutation). For binary sentiment and paraphrase tasks, inversion swaps the two labels (POS ↔ NEG, SIMILAR ↔ DIFFERENT, HATE ↔ NOT_HATE); for three-way NLI, labels cycle as ENTAILMENT → NEUTRAL → CONTRADICTION → ENTAILMENT; for four-way AG News topic classification, labels cycle as (y+1) mod 4.
Predictions come from greedy decoding (temperature 0), generating up to three new tokens and mapping the first generated word to a label via task-specific heuristics; outputs matching no label are marked unk and excluded from metrics. Each test example's demonstrations are randomly sampled without replacement from the training set (excluding the query itself), using fixed seeds so the same demonstration sets are used across all models and conditions. The study uses k ∈ {1, 2, 4, 8} demonstrations and 5 random seeds (0, 1, 2, 42, 123). Inputs are capped at 512 tokens for zero-shot and 2,048 tokens for few-shot prompts, with left-sided padding and truncation. All models were accessed via Hugging Face with identical evaluation pipelines. The value of inverted demonstrations is that they let the author check whether a model can be accurate and prompt-consistent at the same time—something that is trivially true under natural labels, since prompt alignment equals accuracy there, but becomes a real test under inversion.
Why This Matters
Impact on research. The paper reframes the task-learning versus prior-refinement debate with a measurable quantity (the semantic override rate) and reports a hard zero rather than a small number. It extends earlier flipped-label work on large proprietary models into the 1–12B range that dominates open deployment, and it argues the limits of few-shot prompting are semantic rather than computational. It also offers prior alignment as a practical diagnostic: high prior alignment predicts that ICL will succeed.
Real-world applications:
- Choosing adaptation strategies: teams that need non-standard label meanings (for example, domain-specific codes) should budget for symbol tuning, contrastive decoding, or fine-tuning rather than prompt engineering.
- Deployment triage: practitioners can run zero-shot evaluations first, since high prior alignment signals where few-shot prompting will pay off.
- Evaluation design: the alignment decomposition gives an audit method for detecting when a model appears to follow a prompt while actually relying on pre-trained tendencies.
- Benchmark interpretation: results on tasks with weak priors (such as QQP here) can be misleading, because gains may come from generic task signal rather than the demonstrated mapping.
Industry relevance. Any deployment that relies on few-shot prompting with semantically loaded label tokens (POS, NEG, ENTAILMENT) should expect the model to refine its pre-trained behavior rather than adopt new label meanings. Systems requiring arbitrary or inverted label semantics at these model scales need weight updates or inference-time interventions, per the paper's conclusion.
Future Directions
- Test beyond semantically loaded labels. The paper asks whether the constraint is specific to labels with strong meaning or extends to arbitrary symbol-concept mappings.
- Probe the effect of scale more widely. The authors call for exploring how model scale affects the flexibility of these constraints, since only 1B–12B models were tested here.
- Develop the geometric interpretation. The paper hypothesizes that semantic labels occupy topologically stable regions in the representation manifold, and that natural demonstrations move along intrinsic gradients while inverted demonstrations push orthogonally to the manifold; this needs direct testing.
- Compare interventions. Symbol tuning, contrastive decoding, and fine-tuning are cited as possible ways to override semantics, but the paper does not benchmark them against each other or against ICL.
Target Audience
Researchers studying in-context learning mechanisms and the role of pre-training priors; practitioners who deploy small- to mid-sized open-source LLMs (1–12B parameters) using few-shot prompts; and evaluators designing benchmarks who need to distinguish genuine prompt-following from prior-driven behavior. Readers interested in label semantics, prompt sensitivity, and the practical limits of prompt engineering will find the alignment metrics reusable.
Authors’ abstract
Can in-context learning (ICL) override pre-trained label semantics, or does it merely refine an existing semantic backbone? We address this question by treating LLMs as prompt-induced classifiers and contrasting their behavior under \emph{natural} demonstrations (with correct labels) and \emph{inverted} demonstrations (systematically flipping label meanings). We decompose ICL behavior into three alignment metrics (truth, prior, and prompt alignment) and introduce a semantic override rate, defined as correctness under flipped semantics. Across eight classification tasks and eight open-source LLMs (1--12B parameters), we find consistent evidence for a semantic anchor view. With natural demonstrations, ICL improves accuracy while maintaining strong prior alignment; most correct predictions coincide with zero-shot behavior, even when the prior is weak. With inverted demonstrations, models cannot learn coherent anti-semantic classifiers: prompt alignment increases only by sacrificing accuracy, and semantic override rates remain exactly zero in our few-shot 1--12B setting. Rather than flexibly remapping label meanings, ICL primarily adjusts how inputs project onto stable semantic directions learned during pre-training, clarifying fundamental limits of few-shot prompting and suggesting that overriding label semantics at these scales requires interventions beyond ICL. All code is available at: https://github.com/AnanthaPadmanaban-KrishnaKumar/semantic-anchors-icl.