Research
PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection
Overview Research area: Time-series anomaly detection (semi-supervised), with a focus on representation learning rather than forecasting or reconstruction. Technical level: Intermediate. Readers shoul

- arXiv
- 2602.01359
- Published
- 2026-02-01
- Authors
- Jinju Park, Seokho Kang
AI summary
Overview
- Research area: Time-series anomaly detection (semi-supervised), with a focus on representation learning rather than forecasting or reconstruction.
- Technical level: Intermediate. Readers should be comfortable with embedding spaces, metric learning losses, and anomaly detection benchmark metrics.
- Scope: The paper introduces PaAno, a lightweight 1D-CNN, patch-embedding method for univariate and multivariate time-series anomaly detection, and evaluates it on the TSB-AD benchmark against 48 baseline methods.
What This Paper Is About
Recent time-series anomaly detection research has moved toward ever-larger architectures such as Transformers and foundation models, but these are expensive in computation and memory, and prior work (Sarfraz et al. 2024; Liu and Paparrizos 2024) has argued that their reported advantages shrink or disappear under rigorous evaluation protocols. The authors target this gap by asking whether a small, locality-focused model can detect anomalies better and faster. Their goal is a method that learns a discriminative embedding space over short temporal patches and scores anomalies by distance to a memory bank of normal patches.
Key Contributions
- A representation-based framework for anomaly detection. PaAno builds a patch-level embedding space tailored for time-series anomaly detection, addressing what the authors describe as an underexplored direction relative to forecasting- and reconstruction-based methods.
- A lightweight architecture. The method uses a compact 1D-CNN (0.3M parameters in the reported experiments), which the authors position as faster and more efficient than recent heavy neural network architectures.
- State-of-the-art results on both task types. PaAno ranked first across all six performance measures on both univariate (TSB-AD-U) and multivariate (TSB-AD-M) time-series anomaly detection, covering range-wise and point-wise measures.
- Robustness to hyperparameters. The authors report that performance remains stable across different memory bank sizes, numbers of nearest neighbors, patch encoder architectures, loss weights, patch sizes, and minibatch sizes, indicating extensive tuning is unnecessary.
Main Findings
- Best on all measures, both settings: On TSB-AD-U, PaAno scored 0.53 VUS-PR, 0.89 VUS-ROC, 0.49 Range-F1, 0.47 AUC-PR, 0.87 AUC-ROC, and 0.52 Point-F1, ranking 1st on each. On TSB-AD-M, it scored 0.43 VUS-PR, 0.79 VUS-ROC, 0.41 Range-F1, 0.38 AUC-PR, 0.76 AUC-ROC, and 0.43 Point-F1, again ranking 1st on each.
- Efficiency: PaAno used 0.3M parameters with average run times of 5.4s (univariate) and 9.7s (multivariate). For comparison, the paper reports AnomalyTransformer at 4.8M parameters and 48.9s on TSB-AD-U, MOMENT at 109.6M parameters, TimesFM at 203.5M parameters and 83.8s, and Lag-Llama at 1220.8s on TSB-AD-U.
- Baselines were competitive but not better: KAN-AD was second-best on five of six univariate measures (0.40 VUS-PR, 0.82 VUS-ROC, 0.43 Range-F1, 0.41 AUC-PR, 0.80 AUC-ROC, 0.44 Point-F1). (Sub)-PCA recorded the second-best univariate VUS-PR at 0.42. In the multivariate setting, KAN-AD reached 0.40 VUS-PR, while DADA, DeepAnT, OmniAnomaly, and PCA each ranked third at 0.31 VUS-PR.
- Transformer-based methods underperformed despite heavier architectures: The authors state that Transformer-based methods showed relatively low performance in both settings; AnomalyTransformer, for example, scored 0.12 VUS-PR on TSB-AD-U and 0.12 on TSB-AD-M.
- Every component matters: In the ablation study (reported in Table 4, with values expressed as percentages), removing instance normalization dropped VUS-PR to 45.3 on TSB-AD-U and 33.4 on TSB-AD-M; removing both losses dropped it to 48.0 and 35.6; removing the pretext loss alone gave 51.1 and 42.2, versus 53.0 and 42.6 for the full method.
- The specific losses were better than alternatives: Replacing the triplet loss with InfoNCE reduced VUS-PR to 48.3 (univariate) and 36.2 (multivariate). Dropping negative selection in the triplet loss gave 50.9 and 40.2. Using the pretext loss continuously rather than only early degraded results to 47.4 and 40.8 and increased computational cost.
- Online adaptability without retraining: Because the memory bank can be maintained as a queue that inserts recent normal patch embeddings and discards old ones, the authors state the method can track non-stationary normal patterns through memory bank updates alone, without model retraining.
- Evaluation protocol: The study removes point adjustment entirely and adds four threshold-independent measures, following concerns about inflated scores from point adjustment and threshold tuning. VUS-PR is treated as the primary measure.
Methodology in Plain English
The method assumes a training time series containing only normal behavior. It slides a fixed-length window one step at a time over that series to cut out many short overlapping patches. Each patch is standardized internally to zero mean and unit variance, which the authors say makes representations more stable against shifts such as regime changes or drift.
A small 1D-CNN encodes each patch into a vector. Two extra small networks sit on top during training only: an MLP projection head used for metric learning, and a classification head used for a self-supervised task. The training objective combines two losses. The triplet loss pulls a patch closer to a slightly shifted version of itself (a positive) than to the most dissimilar patch in the same minibatch (the farthest negative), by at least a margin. The pretext loss asks the model to judge whether two patches are temporally consecutive, using the patch exactly one window earlier as the positive and random patches as negatives; this loss is used only in the early stage of training and is linearly decayed to zero.
After training, only the patch encoder is kept. The authors pass all training patches through it to build a memory bank of normal embeddings, then shrink that bank with K-means clustering, keeping the vector closest to each of K cluster centroids. At inference, for a query time step, they consider all patches that contain that time step, embed them, and compute each patch's anomaly score as the average cosine distance to its k nearest neighbors in the reduced memory bank. The final score for the time step is the average of those patch-level scores. A high value means the local pattern around that time step does not resemble anything seen in training.
Why This Matters
- Impact on research: The paper argues that locality, not scale, is the lever for this task. It supports the position that evaluation protocol matters more than architecture size, since simpler methods remained competitive and heavy models did not dominate under the TSB-AD protocol without point adjustment.
- Real-world applications:
- Industrial sensor monitoring, where equipment faults appear as spikes, drops, or sustained deviations.
- Financial market transaction monitoring, where unusual patterns need timely flagging.
- Healthcare monitoring, where deviations in patient signals require fast detection.
- Edge or on-device deployments, where memory and compute budgets exclude foundation-model-scale approaches.
- Industry relevance: The 0.3M parameter footprint and 5.4s/9.7s run times make the method plausible for real-time and resource-constrained settings where heavier Transformer and foundation model baselines are impractical. The ability to update the memory bank online without retraining fits streaming production environments where normal behavior drifts.
Future Directions
- Long-term deployment under drift: The paper proposes queue-based memory bank updates as a solution for non-stationary normal regimes, but the experiments reported here do not evaluate a continuously updating bank over a long stream; how scores behave during and after drift remains open.
- Cross-variable structure in multivariate data: PaAno processes patches and relies on instance normalization across channels. The paper does not report an explicit mechanism for capturing the dependencies among variables across time steps that define the multivariate setting, leaving this as a question for further work.
- Extending the pretext task: The pretext loss is used only in early training and is compared against continuous use and against removing the linear decay. Other self-supervised objectives for temporal structure are not explored in the reported results.
- Toward truly sub-second and embedded operation: Reported run times are aggregate averages over benchmark datasets (5.4s and 9.7s), not per-series latencies; whether the method meets the latency demands of specific real-time or embedded applications is not established in the content presented.
Target Audience
Researchers and practitioners in time-series anomaly detection who want a strong, low-cost baseline, particularly those working under semi-supervised settings with only normal training data. It is also useful for engineers deploying detection on constrained hardware, and for researchers interested in evaluation methodology, since the paper frames its claims around the TSB-AD protocol with point adjustment removed. Readers without background in embedding-based anomaly detection will need some familiarity with metric learning losses to follow the training objective.
Authors’ abstract
Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transformers and foundation models, they incur high computational costs and memory usage, making them impractical for real-time and resource-constrained scenarios. Moreover, they often fail to demonstrate significant performance gains over simpler methods under rigorous evaluation protocols. In this study, we propose Patch-based representation learning for time-series Anomaly detection (PaAno), a lightweight yet effective method for fast and efficient time-series anomaly detection. PaAno extracts short temporal patches from time-series training data and uses a 1D convolutional neural network to embed each patch into a vector representation. The model is trained using a combination of triplet loss and pretext loss to ensure the embeddings capture informative temporal patterns from input patches. During inference, the anomaly score at each time step is computed by comparing the embeddings of its surrounding patches to those of normal patches extracted from the training time-series. Evaluated on the TSB-AD benchmark, PaAno achieved state-of-the-art performance, significantly outperforming existing methods, including those based on heavy architectures, on both univariate and multivariate time-series anomaly detection across various range-wise and point-wise performance measures.