Research
Quantized FCA: Efficient Zero-Shot Texture Anomaly Detection
Overview Research area: Computer vision, specifically zero-shot anomaly detection and localization in textures. Technical level: Intermediate. The paper builds on the feature correspondence analysis (

- arXiv
- 2510.15602
- Published
- 2025-10-17
- Authors
- Andrei-Timotei Ardelean, Patrick Rückbeil, Tim Weyrich
AI summary
Overview
Research area: Computer vision, specifically zero-shot anomaly detection and localization in textures.
Technical level: Intermediate. The paper builds on the feature correspondence analysis (FCA) framework and assumes familiarity with anomaly detection metrics such as AUROC and AUPRO, but its central idea (replacing sorting with histograms) is approachable.
Scope: The paper presents QFCA and QFCA+, a real-time reformulation of FCA that speeds up zero-shot texture anomaly localization by an order of magnitude while maintaining or improving accuracy.
What This Paper Is About
Zero-shot anomaly localization (ZSAL) aims to find and outline defective regions in a single image of a texture, without any training or example images of what "normal" or "anomalous" looks like. The best existing method for this task, FCA, is accurate but slow (about 1.07 seconds per image on MVTec AD), which makes it impractical for live use. This paper's goal is to make that method fast enough for real-time deployment while keeping its detection quality.
Key Contributions
-
QFCA, a quantized version of FCA. The paper replaces FCA's expensive per-patch, per-channel sorting operation with a histogram-based comparison over quantized feature values, using a two-pointer algorithm that runs in linear time in the number of bins. This yields a 10× speedup with little to no accuracy loss.
-
A PCA-based feature preprocessing step. The authors add a single-image preprocessing technique that subtracts the PCA reconstruction from the extracted features, reducing the variance of normal features while preserving anomalous ones. The full method using this step is called QFCA+. It improves localization on complex textures at a modest runtime cost (about 40 ms for a 1024×1024 image).
-
A constant-time local average pooling implementation. The authors identify that average pooling in PyTorch, TensorFlow, and JAX scales with kernel size, which makes the overall method depend on patch size. They implement 2D average pooling using summed-area tables (integral images) with O(H × W) complexity per channel, cutting overall running time by 30% in the usual case of patch size 9.
-
A thorough evaluation across three datasets. QFCA and QFCA+ are benchmarked against VLM-based and texture-specific methods on MVTec AD, DTD-Synthetic, and Woven Fabric Textures.
Main Findings
-
QFCA is roughly 10× faster than FCA with comparable accuracy. On MVTec AD, QFCA scores PRO 97.13, AUROC_s 98.72, and F1 71.88 at 0.057 s per image, versus FCA at PRO 97.18, AUROC_s 98.73, and F1 71.75 at 1.070 s per image.
-
QFCA+ gives the best overall metrics on MVTec AD. PRO 97.57, AUROC_s 98.83, and F1 73.07 at 0.097 s per image.
-
Quantization saturates at 16 bins. Using as few as 16 histogram bins is enough to match the non-quantized FCA result; a small gap of 0.05% remains even with many bins, attributed to differences in postprocessing.
-
QFCA+ outperforms FCA and the FCA + k-NN variant on the other datasets. On Woven Fabric Textures it reaches PRO 92.99, AUROC_s 98.51, F1 79.71 at 0.041 s per image (FCA: 89.57 / 98.26 / 79.13 at 0.179 s; FCA + k-NN: 88.77 / 97.73 / 76.30 at 6.345 s). On DTD-Synthetic it reaches PRO 96.64, AUROC_s 98.74, F1 71.57 at 0.029 s per image (FCA + k-NN: 95.93 / 98.51 / 71.79 at 3.970 s).
-
QFCA clearly beats VLM-based zero-shot methods on textures. On MVTec AD, WinCLIP scores PRO 71.50, AUROC_s 89.06, F1 38.42 at 0.389 s; SAA+ scores 64.79 / 77.82 / 59.19 at 0.270 s; April-GAN scores 92.56 / 96.51 / 58.63 at 0.122 s; SDP+ scores 92.48 / 96.76 / 52.70 at 0.045 s.
-
Running QFCA at reduced resolution is a practical option. QFCA at 512×512 scores PRO 97.08, AUROC_s 98.77, F1 68.91 in 0.014 s per image.
-
Median-based reference selection works best. Comparing reference procedures, median followed by quantization gave PRO 97.1, AUROC_s 98.7, F1 71.9; quantize-then-mean gave 96.9 / 98.6 / 71.3.
-
The spatial weighting Gaussian peaks at a simple average. Accuracy increases with σ_p and peaks when σ_p = ∞, corresponding to simple average pooling, meaning all patches a pixel belongs to are treated as equally important. Default final smoothing is σ_s = 1.
-
Optimal patch size scales with image size. Empirically, the optimal patch size is about 10% of the feature map size, or 1.25% of the pre-feature-extraction image size. Patch sizes tested were 3, 5, 7, 9, and 11.
-
10 principal components is a robust default for the preprocessing step; the optimum depends on image complexity.
-
On generic (non-texture) objects, the method is weaker. Across all MVTec AD objects, QFCA+ averages PRO 74.18, AUROC_s 84.48, F1 40.69, AUROC_c 74.45, compared with FCA at 68.78 / 83.70 / 38.76 / 67.23. The authors note it remains on par with most VLM-based methods despite being texture-specific.
Methodology in Plain English
The starting point is FCA, which describes an anomaly score for a pixel as the sum of errors obtained by comparing the local patch around that pixel with a global reference that characterizes the texture. FCA computes this comparison by sorting feature values in each patch and in the reference, then matching values of the same rank and measuring their differences. That sorting, repeated for every patch and every channel, is the bottleneck.
QFCA replaces the sorting with quantization. Feature values are mapped into N equally spaced bins, so each patch becomes a histogram and the reference becomes another histogram. A two-pointer sweep walks through both histograms, moving mass between matching quantiles and accumulating the weighted absolute difference between bin values. The sweep finishes in 2N − 1 steps, so cost grows linearly with the number of bins rather than with the square of the patch size times its logarithm as in FCA. The researchers compile this into a CUDA kernel so all patches and channels are processed in parallel.
Error-to-pixel association is preserved from FCA: after computing a per-bin error, the error is distributed back to all pixels that fell in that bin, aggregated using a Gaussian blur whose kernel size equals the patch size. Because this blur and the histogram formation are local average pooling operations, the authors also replace the framework pooling calls with a summed-area-table implementation that is independent of kernel size.
Reference selection follows FCA's median approach: the reference is the distribution minimizing the Wasserstein distance across all patches, computed in full precision and quantized afterward. For the QFCA+ variant, the feature extractor (Wide ResNet-50 features taken after the second convolutional block) is augmented by a preprocessing step: PCA is applied to the features and the reconstruction is subtracted from them, so that features well explained by the dominant principal components are suppressed and rare, anomalous features stand out. This is a single-image operation, intended as a cheaper alternative to FCA's k-NN variant, which scales quadratically with the number of patches.
Evaluation uses three datasets: MVTec AD (five texture classes: carpet, grid, leather, tile, wood, over 500 images with masks), DTD-Synthetic (12 textures, 100 training images and over 100 test images each, sizes from 180×180 to 384×384), and Woven Fabric Textures (2 textures, 50 images each, 512×512 masks with ground truth). Metrics are pixel-level AUROC_s, AUPRO computed up to a false positive rate of 30% (called PRO in the paper), F1 at the optimal threshold, and per-image latency.
Why This Matters
This work matters because it moves zero-shot texture anomaly localization from an offline research setting into a real-time one. Prior art was accurate but too slow to be useful for live monitoring; QFCA+ runs in tens of milliseconds per image. The paper also contributes a general-purpose speedup (summed-area-table average pooling) that could benefit other pipelines, and a feature preprocessing trick that can be attached to anomaly detection methods beyond FCA.
Real-world applications named in the paper:
- Manufacturing inspection and assembly line monitoring, where live feeds must be checked for surface defects.
- Medical imaging, where anomalous regions must be segmented from otherwise homogeneous data.
- Machine vision and data analysis preprocessing, where outliers must be flagged before further processing.
- Other signal domains the paper cites for anomaly detection broadly, including acoustic and video monitoring, weather records, and financial fraud detection.
Industry relevance: The paper explicitly frames assembly line monitoring as the motivating deployment scenario and reports latency rather than throughput to reflect a live-feed use case. The ability to run at interactive rates, and the option of a lower-resolution mode (QFCA 512 at 0.014 s per image) or a higher-accuracy mode (QFCA+ at 0.097 s per image), gives practitioners a tunable accuracy-versus-speed tradeoff on the same algorithm.
Future Directions
-
Extending beyond textures. The authors state as a limitation that, without a large VLM injecting general knowledge, QFCA is suitable for texture-like data and not for arbitrary objects. Closing that gap is the natural next step.
-
Choosing the number of principal components automatically. The paper reports that the optimal number of components depends on image complexity, with 10 working as a robust off-the-shelf choice, which suggests an adaptive selection rule as an open problem.
-
Reusing the fast pooling primitive. The summed-area-table average pooling implementation is presented as potentially includable in other pipelines; verifying and integrating it more broadly, and reporting its adoption upstream in PyTorch, TensorFlow, or JAX, is an open direction.
-
Reconciling the residual 0.05% gap with FCA. The paper attributes the small remaining difference to postprocessing rather than quantization, which invites a closer analysis of the postprocessing pipeline.
Target Audience
Researchers and engineers working on visual anomaly detection, particularly those interested in zero-shot or unsupervised methods for textures. It is also relevant to practitioners who need deployable inspection systems and care about the accuracy-latency tradeoff, and to anyone working on efficient GPU implementations of statistical comparison operations. Readers should be comfortable with anomaly detection metrics (AUROC, PRO/AUPRO, F1) and with concepts such as PCA and Wasserstein distance, though the paper keeps the algorithmic exposition largely self-contained.
Authors’ abstract
Zero-shot anomaly localization is a rising field in computer vision research, with important progress in recent years. This work focuses on the problem of detecting and localizing anomalies in textures, where anomalies can be defined as the regions that deviate from the overall statistics, violating the stationarity assumption. The main limitation of existing methods is their high running time, making them impractical for deployment in real-world scenarios, such as assembly line monitoring. We propose a real-time method, named QFCA, which implements a quantized version of the feature correspondence analysis (FCA) algorithm. By carefully adapting the patch statistics comparison to work on histograms of quantized values, we obtain a 10x speedup with little to no loss in accuracy. Moreover, we introduce a feature preprocessing step based on principal component analysis, which enhances the contrast between normal and anomalous features, improving the detection precision on complex textures. Our method is thoroughly evaluated against prior art, comparing favorably with existing methods. Project page: https://reality.tf.fau.de/pub/ardelean2025quantized.html