Skip to content
AI.info

Research

FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification

FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification Overview Research area: Efficient machine learning and algorithmic fairness in medical image analysis, specifica

FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification
arXiv
2602.23192
Published
2026-02-26
Authors
Thomas Woergaard, Raghavendra Selvan

AI summary

FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification

Overview

Research area: Efficient machine learning and algorithmic fairness in medical image analysis, specifically mixed-precision quantization for dermatology classifiers.

Technical level: Intermediate. The paper combines quantization theory (uniform quantization, straight-through estimators, learnable bit-widths) with group-fairness objectives, but the core ideas can be grasped without deep quantization expertise.

Scope: The paper proposes FairQuant, a framework that decides which parts of a neural network should keep higher numerical precision and which can be compressed more aggressively, so that compressed dermatology classifiers remain reliable for the worst-served patient subgroup under a fixed bit budget.

What This Paper Is About

Quantizing a neural network (storing its weights at low precision) saves memory and computation, but existing quantization methods only try to preserve average accuracy. In medical settings, average accuracy can hide large performance gaps across patient groups, such as lighter versus darker skin types. The goal of this work is to allocate limited precision across the network in a way that accounts for how much each patient group depends on each part of the model, improving worst-group performance under an explicit bit budget.

Key Contributions

  1. A group-aware importance analysis method. FairQuant runs a short calibration stage (50 mini-batches) in which gradient passes are restricted to one sensitive group at a time, producing per-group sensitivity scores for each quantization scope (per-tensor or per-channel). These are normalized per group and reduced with a max across groups into a single importance map per layer.

  2. A budgeted mixed-precision allocation rule. Importance values are pooled across all scopes and tiered into discrete bit-widths using empirical quantile thresholds derived from a chosen bit palette and target proportions, producing a fixed assignment pattern under a global bit budget.

  3. Bit-Aware Quantization (BAQ). Bit-widths become trainable parameters: each scope carries a real logit that maps through a tanh to a continuous bit proxy within a fixed interval, rounded with a straight-through estimator. Training jointly optimizes weights, bit logits, a fairness penalty, and an L2 regularizer on the logits that pushes precision toward the minimum bit value.

  4. An empirical study across backbones and datasets. Evaluation on two dermatology benchmarks with two convolutional backbones (ResNet18/50) and two compact vision transformers (DeiT-Tiny, TinyViT), reporting average accuracy, worst-group accuracy, equalised opportunity gap (EOpp0), equalized odds gap (EOdd), effective bits per parameter, and giga bit-operations (GBOPs), plus stability ablations.

Main Findings

  • Uniform 8-bit tracks full precision; uniform 4-bit can fail badly. On Fitzpatrick17k, ResNet18 drops from AvgAcc 50.6 (FP32) to 23.4 with Uniform 4-bit, and from WorstAcc 44.1 to 19.0. TinyViT collapses to AvgAcc 3.0 and WorstAcc 2.0 under Uniform 4-bit on the same dataset, while Uniform 8-bit stays close to FP32.

  • FairQuant recovers much of the 8-bit accuracy at roughly 4 average bits. On Fitzpatrick17k ResNet18, FairQuant with BAQ (FQ-BAQ) at 4.07 average bits reaches AvgAcc 45.33 and WorstAcc 41.53. On TinyViT it reaches AvgAcc 53.60 and WorstAcc 48.17 at 4.12 average bits, approaching the FP32/U8 region; on DeiT-Tiny it reaches AvgAcc 51.83 and WorstAcc 46.00 at 4.30 average bits.

  • The same pattern holds on ISIC 2019. Uniform 4-bit collapses for TinyViT (AvgAcc 53.3, WorstAcc 47.6) and ResNet50 (AvgAcc 57.8, WorstAcc 51.9), and for DeiT-Tiny it preserves AvgAcc 81.4 but lowers WorstAcc to 79.1 with a gap of 5.7. FQ-BAQ at roughly 4.1–4.4 average bits reaches AvgAcc 82.73 / WorstAcc 81.13 (TinyViT) and AvgAcc 82.17 / WorstAcc 80.83 (ResNet50).

  • Fixed mixed-precision QAT (FQ-QAT) often gives the highest accuracies at slightly higher average precision. ISIC DeiT-Tiny FQ-QAT reaches AvgAcc 83.80 and WorstAcc 82.87 at 5.04 average bits; ISIC TinyViT reaches 83.40 / 82.47 at 4.38 average bits; ISIC ResNet50 reaches 84.87 / 82.67 at 4.39 average bits.

  • Comparison to prior fairness-aware compression is mixed but generally favourable. FairGRAPE (Lin et al.) at 4.80 average bits reports AvgAcc 32.60 / WorstAcc 28.70 on Fitzpatrick17k DeiT-Tiny and 13.40 / 12.00 on TinyViT, and Guo et al. at 4.92 bits report AvgAcc 40.80 / WorstAcc 36.60 on Fitzpatrick17k ResNet18. On ISIC ResNet50, Lin et al. at 4.80 average bits reports a higher AvgAcc of 86.20 with WorstAcc 79.70, which is above FQ-QAT and FQ-BAQ on average accuracy but below both on worst-group accuracy.

  • The method is stable under its main hyperparameters. Sweeping the BAQ bitrate regularizer λ_baq,b on ResNet18 across both datasets, AvgAcc and WorstAcc vary only modestly with tight bands across seeds, and EOdd stays within a narrow band; the default λ_baq,b = 0.01 is chosen from this flat region.

  • Fairness regularization trades accuracy for gap closure. Increasing λ_fair generally improves EOdd, most clearly at the high end of the sweep, but with a noticeable drop in both AvgAcc and WorstAcc. The paper notes that part of this apparent EOdd improvement may reflect a less accurate classifier whose errors spread across groups, narrowing the measured disparity rather than producing a strictly better operating point. The default λ_fair = 0.5 is selected from a moderate regime.

  • BAQ is robust to learning-rate choice in the low range. Performance is stable at low learning rates for the bit proxies, then degrades sharply at the largest values, with EOdd shifting at the same time.

  • Pretraining amount differs across backbones. ResNet models are pretrained for 200 epochs; DeiT-Tiny and TinyViT for 20 epochs, with all QAT and FairQuant runs fine-tuned for 10 epochs using AdamW, batch size 128, weight decay 0.01, and a learning rate of 10⁻⁴ to epoch 160 followed by a ×0.1 decay.

  • An acknowledged confound on ISIC. The pretrained ResNet50 baseline and Uniform quantization have lower average accuracy than the QAT and BAQ variants on ISIC, which the authors attribute to the additional training performed during QAT and BAQ, noting a similar increase is reported in FairQuantize (Guo et al.).

Methodology in Plain English

The pipeline has three stages. First, the frozen full-precision model is run over a small calibration set drawn from the training split, with backward passes restricted to one sensitive group at a time, so the authors can see how strongly each group depends on each layer or channel. Second, these per-group sensitivity scores are normalized and combined into a single importance map, and the map is sliced into tiers by quantiles so that a fixed percentage of scopes get each bit-width from a chosen palette (the main experiments use the allocation {20%: 2, 40%: 4, 40%: 8} for QAT and a 4–8 range for BAQ). Third, in the BAQ variant, each scope's bit-width is replaced by a trainable value bounded inside an interval; the model is trained on the task loss plus a fairness term equal to the difference between the largest and smallest group loss, plus a penalty that shrinks the bit logits and pushes precision toward the minimum. Weights and bit-widths are learned at the same time, gradients flow through the rounding step via a straight-through estimator, and the rounded bit-widths are fixed at the end for inference. Comparisons use FP32, Uniform-8, Uniform-4, FairGRAPE, and Guo et al. as baselines; main results use three seeds per setting and ablations use five seeds with 95% confidence intervals, all on a ResNet18 backbone.

Why This Matters

The work reframes precision allocation as a fairness decision rather than a purely engineering one, showing that the standard practice of tuning compression for average accuracy can silently damage the worst-served patient group. It contributes a reusable recipe: group-conditioned importance signals, an explicit budget rule, and a trainable allocation scheme that can be dropped into existing quantization-aware training.

Real-world applications:

  • Teledermatology triage tools that must run on limited hardware while remaining reliable across skin tones.
  • Smartphone-based skin screening applications where memory and battery budgets constrain model size.
  • Clinical decision support deployed in clinics with constrained compute, latency, or energy budgets.
  • Any medical imaging pipeline where a sensitive attribute (such as sex in ISIC 2019) must not be systematically disadvantaged by model compression.

Industry relevance: Practitioners deploying medical imaging models on edge devices need to meet hard latency and memory budgets. FairQuant offers a way to select precision per layer with explicit control over the bit budget and an adjustable fairness pressure, rather than choosing a single global bit-width by trial and error. The released code at https://github.com/saintslab/FairQuant lowers the barrier to reproducing and extending the approach.

Future Directions

  1. Broaden the fairness sweep. The authors state they only ran a coarse sweep over λ_fair and that a denser, better targeted sweep could reveal a clearer picture of its effect and identify a stronger operating point.

  2. Move beyond a batch-level fairness proxy. The current fairness term approximates group disparities within a mini-batch and is evaluated with EOpp and EOdd; the authors note other fairness definitions can change which operating point is preferred.

  3. Extend to more datasets, attributes, and label quality regimes. The evaluation covers two dermatology datasets and two sensitive attributes (skin type and sex), and the paper notes that group labels can be noisy or incomplete.

  4. Test generalization of the learned allocation. The paper's own framing asks whether a single approach can improve subgroup outcomes under a strict bit budget without tailoring each model manually, and whether the learned allocation stays stable across seeds and training noise — questions worth probing across additional backbones, budgets, and deployment targets.

Target Audience

Researchers and practitioners working at the intersection of model compression and algorithmic fairness, particularly those applying quantization to medical imaging. It is also relevant to machine learning engineers who need to deploy dermatology or other clinical classifiers under strict memory, latency, and energy budgets, and to fairness researchers interested in how compression choices propagate into subgroup performance.

Authors’ abstract

Compressing neural networks by quantizing model parameters offers useful trade-off between performance and efficiency. Methods like quantization-aware training and post-training quantization strive to maintain the downstream performance of compressed models compared to the full precision models. However, these techniques do not explicitly consider the impact on algorithmic fairness. In this work, we study fairness-aware mixed-precision quantization schemes for medical image classification under explicit bit budgets. We introduce FairQuant, a framework that combines group-aware importance analysis, budgeted mixed-precision allocation, and a learnable Bit-Aware Quantization (BAQ) mode that jointly optimizes weights and per-unit bit allocations under bitrate and fairness regularization. We evaluate the method on Fitzpatrick17k and ISIC2019 across ResNet18/50, DeiT-Tiny, and TinyViT. Results show that FairQuant configurations with average precision near 4-6 bits recover much of the Uniform 8-bit accuracy while improving worst-group performance relative to Uniform 4- and 8-bit baselines, with comparable fairness metrics under shared budgets.

Read the original paper