Skip to content
AI.info

Research

AP-OOD: Attention Pooling for Out-of-Distribution Detection

AP-OOD: Attention Pooling for Out-of-Distribution Detection Overview Research area: Out-of-distribution (OOD) detection for natural language (and, secondarily, audio), specifically for autoregressive

AP-OOD: Attention Pooling for Out-of-Distribution Detection
arXiv
2602.06031
Published
2026-02-05
Authors
Claus Hofmann, Christian Huber, Bernhard Lehner, Daniel Klotz, Sepp Hochreiter, Werner Zellinger

AI summary

AP-OOD: Attention Pooling for Out-of-Distribution Detection

Overview

Research area: Out-of-distribution (OOD) detection for natural language (and, secondarily, audio), specifically for autoregressive language models such as summarization and translation systems. The work sits at the intersection of uncertainty estimation, representation learning, and semi-supervised anomaly detection.

Technical level: Advanced. The paper builds on the Mahalanobis distance, derives a directional decomposition, and recasts it through attention pooling with a matrix-valued softmax. Readers will need familiarity with transformer encoders/decoders, embedding spaces, and OOD evaluation metrics (AUROC, FPR95).

One-sentence scope: The paper proposes AP-OOD, a semi-supervised OOD detector that replaces mean pooling of token embeddings with learnable attention pooling inside a Mahalanobis-style distance, and demonstrates large FPR95 reductions on text summarization and translation as well as an audio defect-detection task.

What This Paper Is About

Existing OOD detectors for language models typically compress a whole sequence of token embeddings into a single average vector and then measure its distance from a fitted distribution. Averaging discards token-level structure, so sequences built from unusual token combinations can be indistinguishable from in-distribution (ID) sequences once averaged (the paper illustrates this with a toy example where ID and OOD sequence means both cluster at the origin). AP-OOD's goal is to keep that token-level information by learning which tokens to attend to, and to do so in a way that works with no outlier data at all and improves smoothly as small amounts of auxiliary outlier data become available.

Key Contributions

  1. AP-OOD, a token-aware OOD detector for natural language. The method generalizes the Mahalanobis-distance decomposition into directional weight vectors by replacing mean pooling with attention pooling, where each weight vector acts as a learnable query that selects which tokens dominate the pooled representation.

  2. A semi-supervised formulation that interpolates between unsupervised and supervised operation. With no auxiliary (AUX) outlier data the method trains on ID sequences alone; with AUX data it adds a binary cross-entropy term that pushes AUX samples away from the learned prototypes, controlled by a coefficient λ (λ = 0 recovers the unsupervised loss).

  3. Demonstrated state-of-the-art OOD detection for text summarization and translation (and competitive results on an industrial audio dataset), including improvements over mean-pooling baselines on both input token embeddings and decoder output embeddings.

  4. A theoretical motivation for why the approach suits tokenized data. The authors show that when the inverse temperature β = 0 and the number of heads M equals the embedding dimension D, their distance reduces to the Mahalanobis distance; they also derive a min-based score and show the score actually used is its upper bound, and analyze the extended distance through kernel functions.

Main Findings

  • Large FPR95 gains on summarization. On XSUM summarization with PEGASUS LARGE, the abstract reports that AP-OOD reduces FPR95 from 27.84% to 4.67%. In Table 1 (input OOD, averaged over the CNN/DM, Newsroom, Reddit, and Samsum OOD sets), the best baseline is Deep SVDD at 27.84% mean FPR95, while AP-OOD reaches 4.67% mean FPR95.

  • Per-dataset input results. For CNN/DM input OOD, FPR95 improves from 74.20% (best baseline Deep SVDD) to 12.88% (AP-OOD). AP-OOD's mean input AUROC across the four OOD sets is 98.91%, versus 91.60% for the next-best baseline listed. On the Samsum input set, AP-OOD reaches 99.77% AUROC with FPR95 of 0.00%.

  • Output-side gains. Averaged over the evaluated OOD sets, AP-OOD's output-side mean FPR95 is 16.26% with mean AUROC 96.39%; the best-baseline mean FPR95 listed is Deep SVDD at 32.33%, and the best-baseline mean AUROC listed is binary logits at 88.49%.

  • Embedding-based methods beat prediction-based ones on summarization. The table shows Mahalanobis, KNN, Deep SVDD, and AP-OOD outperforming perplexity and entropy on the summarization task.

  • Translation improvements are reported but smaller. The abstract states FPR95 falls from 77.08% to 70.37% on WMT15 En–Fr. In the portion of Table 2 visible in the supplied content, AP-OOD attains the best mean input AUROC (74.81%) and the best mean input FPR95 (75.25%) among the unsupervised methods listed, with the best value in every individual input OOD column. The remainder of Table 2 is truncated in the provided text, so the complete output-side numbers are not available here.

  • Task-dependent competitiveness of perplexity. Unlike in summarization, prediction-based methods (perplexity and entropy) are competitive with embedding-based methods in translation. The authors attribute this to differing aleatoric uncertainty: translation constrains output lexically and syntactically (low-entropy outputs), whereas summarization admits multiple valid outputs (higher output entropy).

  • AUX data helps, and more AUX data helps more. In the semi-supervised experiments on CNN/DM and Newsroom input embeddings (Figure 3), supplying AUX data improves AUROC for AP-OOD, and AP-OOD achieves the highest AUROC independent of the AUX sample count, compared against binary logits, Deep SAD, and relative Mahalanobis.

  • Scaling helps. Increasing the number of heads M and queries T produces a steady rise in mean AUROC on summarization; the largest configuration tested (M = 1024, T = 16) reaches 99.40% mean AUROC, in a setting where the largest configuration has 16 times the parameter count of the Mahalanobis baseline.

  • Audio results. On the MIMII-DG audio dataset the underlying classifier reaches 97.6% accuracy on the primary classification task, and AP-OOD improves FPR95 from 36.43% (MSP) to 22.35%. (The full audio results table, Table 3, is not visible in the provided excerpt.)

  • Runtime is acceptable. AP-OOD is slower than the Mahalanobis baseline but substantially faster than a forward pass through the transformer encoder, and it scales linearly, so its relative overhead shrinks for longer sequences.

  • Umbrella score choice. Comparing the min-based score s_min to its upper-bound variant s, the authors find empirically (Appendix D.7) that the upper-bound score gives stronger OOD discrimination.

Methodology in Plain English

The starting point is the Mahalanobis distance: fit a Gaussian to the per-sequence mean embeddings of the training corpus, then score a new sequence by how far its mean embedding is from that Gaussian's center. The authors first rewrite this distance as a sum of squared projections onto a set of weight vectors that span the embedding space. This decomposition makes the assumption behind mean pooling visible: every token contributes equally, and anomalies in individual tokens get averaged away.

AP-OOD replaces averaging with softmax attention. Each weight vector becomes a learned query; the query scores every token in the sequence, the scores are turned into a softmax distribution (modulated by an inverse temperature β), and the sequence is summarized as a weighted combination of its tokens. Instead of one fixed prototype, AP-OOD also builds a corpus-wide prototype the same way, by concatenating many sequence representations and attention-pooling over them. A sequence's score is the negative squared distance between its attention-pooled representation and the corpus prototype, summed over M heads, plus a log-norm regularizer on the queries. Training minimizes within-ID squared distances while the log term prevents the query norms from collapsing.

The authors extend the scheme so each head can carry T queries, producing a matrix of pooled representations rather than a vector, using a matrix-valued softmax that normalizes over both rows and columns; the final scalar is a Frobenius inner product with the query matrix. They also provide a mini-batch version of the corpus-wide pooling so that the concatenated representation does not have to be held in memory.

For the supervised variant, the authors follow the outlier-exposure idea: given N ID sequences and N′ auxiliary outlier sequences, they minimize the ID squared distances and add a term that penalizes small distances for AUX samples, weighted by λ. Because λ = 0 reduces to the unsupervised loss, the method slides continuously between the two regimes as outlier data accumulates. Across experiments the transformer backbone stays frozen and only the AP-OOD parameters are trained.

The empirical evaluation covers three settings: text summarization with PEGASUS LARGE fine-tuned on XSUM (C4 used as AUX; CNN/Daily Mail, Newsroom, Reddit TIFU, and Samsum as OOD sets; ForumSum excluded because the dataset used in prior work was retracted), English-to-French translation with a base Transformer on WMT15 En–Fr (ParaCrawl En–Fr as AUX; newstest2014, newsdiscussdev2015, newsdiscusstest2015, and OPUS Law/Koran/Medical/IT/Subtitles as OOD), and audio using MIMII-DG. Training used 100,000 ID sequence representations in all experiments, Adam with learning rate 0.01 and no weight decay, a cosine schedule, 1,000 steps, and batch size 512; M and T were selected to match the Mahalanobis baseline's parameter count. The translation Transformer trained for 100,000 steps with AdamW, a cosine schedule, linear warmup, peak learning rate 5×10⁻⁴, batch size 1024, and context length 512. Reported standard deviations come from five independent data-set splits and training runs, and evaluation uses FPR95 and AUROC.

Why This Matters

Impact on research. The paper challenges a default assumption in embedding-based OOD detection — that a sequence should be summarized by its mean. By showing that the Mahalanobis distance can be decomposed into learnable directional queries and then re-pooled with attention, it gives the community a concrete, drop-in upgrade path for existing distance-based detectors. The semi-supervised interpolation also addresses a practical complaint about outlier exposure: large, diverse AUX corpora are often unobtainable, so a detector that degrades gracefully with few outliers is more usable.

Real-world applications:

  • Abstractive summarization systems that must recognize when a prompt falls outside their training domain — the paper cites the failure case where a model trained on BBC articles outputs "All images are copyrighted" for CNN articles.
  • Machine translation services that need to flag source sentences from unseen domains (legal, medical, religious, subtitles) before emitting low-quality output.
  • Hallucination mitigation, since the paper links OOD prompts to high epistemic uncertainty and positions detection as a way to notify users that no reliable output can be generated.
  • Industrial monitoring, specifically defect detection from machine audio recordings, where ID recordings are plentiful but defective-machine recordings are rare and hard to collect.

Industry relevance. The method is post-hoc and keeps the base model frozen, so it can be added on top of an already-deployed model. It scales linearly, adds less overhead than a transformer forward pass, and its supervision knob lets teams start unsupervised and improve the detector as they accumulate domain-specific outlier data. The authors release code at https://github.com/ml-jku/ap-ood.

Future Directions

  • Extending the evaluation beyond the three settings shown. The paper verifies the method on decoder-only language modeling with Pythia-160M in an appendix, but broader coverage of model families, languages, and long-context tasks is left open.

  • Determining how much AUX data is enough. The semi-supervised results (Figure 3) show monotone improvement with more AUX samples up to 10,000, but the question of selecting or curating a small, highly informative AUX set rather than a large generic one is not resolved.

  • Choosing M, T, and β in a principled way. The scaling study finds 99.40% mean AUROC at M = 1024, T = 16, the largest configuration tested, raising the question of whether further scaling continues to help and how practitioners should pick these hyperparameters without an OOD validation set.

  • Understanding when prediction-based scores are competitive. The paper's explanation for perplexity's strength on translation versus summarization is framed around aleatoric uncertainty; turning that into a predictive guideline for which family of detectors to use on a new task remains an open question.

Target Audience

This paper is best suited to machine learning researchers and practitioners working on uncertainty estimation, anomaly detection, or the reliability of deployed language models, particularly those handling summarization and translation. It will also interest engineers who need a post-hoc OOD layer that can be bolted onto a frozen pretrained model and tuned with whatever outlier data happens to be available. Readers should be comfortable with Mahalanobis distances, attention mechanisms, and OOD evaluation metrics; the derivations in the appendices are aimed at readers who want to verify the connection between the proposed score and the Mahalanobis special case.

Authors’ abstract

Out-of-distribution (OOD) detection, which maps high-dimensional data into a scalar OOD score, is critical for the reliable deployment of machine learning models. A key challenge in recent research is how to effectively leverage and aggregate token embeddings from language models to obtain the OOD score. In this work, we propose AP-OOD, a novel OOD detection method for natural language that goes beyond simple average-based aggregation by exploiting token-level information. AP-OOD is a semi-supervised approach that flexibly interpolates between unsupervised and supervised settings, enabling the use of limited auxiliary outlier data. Empirically, AP-OOD sets a new state of the art in OOD detection for text: in the unsupervised setting, it reduces the FPR95 (false positive rate at 95% true positives) from 27.84% to 4.67% on XSUM summarization, and from 77.08% to 70.37% on WMT15 En-Fr translation.

Read the original paper