Research
L2CU: Learning to Complement Unseen Users
Overview Research area: Human-AI cooperation in machine learning, specifically the "learning to complement" (L2C) paradigm for classification tasks where a model must work alongside human annotators.

- arXiv
- 2601.06119
- Published
- 2026-01-03
- Authors
- Dileepa Pitawela, Gustavo Carneiro, Hsiang-Ting Chen
AI summary
Overview
Research area: Human-AI cooperation in machine learning, specifically the "learning to complement" (L2C) paradigm for classification tasks where a model must work alongside human annotators.
Technical level: Intermediate. Readers should be comfortable with classification benchmarks, clustering, confusion/transition matrices, and the general distinction between learning-to-defer (L2D) and learning-to-complement (L2C).
Scope: The paper introduces L2CU, a model-agnostic framework that identifies clusters of annotators ("annotator profiles"), trains a cooperative model per profile, and at test time matches a previously unseen user to one of those profiles so the matching model can correct that user's characteristic labeling errors.
What This Paper Is About
Human-AI cooperation systems that "learn to complement" a person's strengths usually assume they will work with the same people they were trained with. The paper targets the harder case: a model trained on a sparse, noisy, multi-annotator dataset that must then cooperate with a user it has never seen. Existing L2C methods represent users with a single global model of human behavior, which ignores that different people make different kinds of mistakes, so cooperative accuracy suffers for unseen users.
Key Contributions
- L2CU framework. A new learning-to-complement framework explicitly designed to complement users unseen during training, built on identifying representative annotator profiles and training one cooperative model per profile.
- Sparse multi-user handling with noisy-label augmentation. L2CU operates when each training annotator labels only a subset of samples, and adds a label augmentation method that generates additional noisy labels from each profile's estimated label transition matrix while preserving that profile's characteristic error pattern.
- Profile matching at test time. A user profiling procedure in which an unseen user annotates a small validation set of M samples per class, a one-versus-all SVM assigns the user to a profile, and an entry condition decides whether the cooperative model should be used at all.
- A new evaluation metric, alteration rate. Positive alteration (A+) and negative alteration (A−) quantify how often the model corrects a user's wrong labels versus how often it corrupts a user's correct labels, alongside original and post-alteration accuracy.
Main Findings
- L2CU leads on every dataset tested. In the accuracy comparison table, L2CU reaches 0.968 ± 0.002 on CIFAR-10, 0.989 ± 0.001 on CIFAR-10N, 0.993 ± 0.002 on CIFAR-10H, 0.878 ± 0.008 on Fashion-MNIST-H, and 0.991 ± 0.004 on Chaoyang. Comparable numbers from competing methods include LECODU at 0.951 (CIFAR-10), 0.989 (CIFAR-10H) and 0.990 (Chaoyang); LECOMH at 0.988 (CIFAR-10H) and 0.988 (Chaoyang); WSP at 0.976 (CIFAR-10H) and 0.872 (Chaoyang); and L2D-Pop at 0.947 (CIFAR-10) and 0.970 (Chaoyang).
- Every profiled unseen test user improved. In the main results table, all test users who were profiled and met the entry condition were classified as improved (I), with zero maintained (M) and zero not-improved (NI): 15/15 on CIFAR-10, 80/80 on CIFAR-10N, 2022/2022 on CIFAR-10H, 182/182 on Fashion-MNIST-H, and 2/2 on Chaoyang.
- Post-alteration accuracy gains. Reported approximate improvements over users' original accuracy are 18% (CIFAR-10N), 5% (CIFAR-10H), 32% (Fashion-MNIST-H) and 15% (Chaoyang). Original versus post-alteration accuracy: 0.836 → 0.989 (CIFAR-10N), 0.940 → 0.993 (CIFAR-10H), 0.662 → 0.878 (F-MNIST-H), 0.858 → 0.988 (Chaoyang), and 0.880 → 0.968 in the CIFAR-10 simulation.
- Positive alterations dominate negative ones. A+ versus A− values are 0.953 vs. 0.09 (CIFAR-10 simulation), 0.954 vs. 0.004 (CIFAR-10N), 0.939 vs. 0.004 (CIFAR-10H), 0.758 vs. 0.073 (F-MNIST-H) and 0.968 vs. 0.045 (Chaoyang).
- Profiles are essential. Removing annotator profiles degrades results across all datasets: CIFAR-10 drops to 5 improved / 10 not improved with post-alteration accuracy 0.835, CIFAR-10N to 69 improved / 11 not improved at 0.918, CIFAR-10H to 1949 improved / 73 not improved at 0.911, F-MNIST-H to 166 improved / 2 maintained / 14 not improved at 0.854, and Chaoyang to 1 improved / 1 not improved at 0.915.
- Cooperation can fix cases where both parties are wrong. The joint-decision distribution shows a non-zero ✗/✗/✓ rate: 00.05% on CIFAR-10N, 00.05% on CIFAR-10H, 04.29% on Fashion-MNIST-H, and 00.13% on Chaoyang. The paper's stated necessary condition is P(C | ¬A, ¬B) > 0.
- Works in the text domain. On an AgNews simulation with K = 3 (silhouette score 0.44), all 15 test users improved, original accuracy 0.700 rose to post-alteration 0.980, with A+ = 0.975 and A− = 0.016.
- Component ablation. With the human label encoder and decision model both removed, post-alteration accuracy is 0.705 (A+ 0.007, A− 0.159); encoder only 0.774 (0.043, 0.083); decision model only 0.861 (0.833, 0.134); both present 0.989 (0.954, 0.004).
- Augmentation count ablation (G). Accuracy at G = 0 is 0.615 (A+ 0.411, A− 0.302), jumping to 0.980 at G = 1, 0.983 at G = 3, and 0.989 at G = 5, with A− stable at 0.004.
- Noise-rate robustness. Across asymmetric noise rates, post-alteration accuracy is 0.992 (40%), 0.968 (60%), 0.879 (80%) and 0.868 (90%), which the paper describes as remaining above 86%.
- Backbone agnostic. Post-alteration accuracy is 0.968 for ResNet-50, 0.969 for DenseNet-121, and 0.989 for ViT/B-16.
- Hard assignment beats soft assignment. Assigning each user to the single best-matching profile model outperformed averaging predictions across all profile models weighted by SVM probabilities.
- Silhouette scores are low on large real annotator pools. CIFAR-10N and CIFAR-10H both show K = 2 with a silhouette score of 0.01, Fashion-MNIST-H K = 2 at 0.09, CIFAR-10 simulation K = 3 at 0.34, and Chaoyang K = 3 at 0.99, which the paper attributes to many annotators each contributing subtle noise patterns.
Methodology in Plain English
Training. Starting from a dataset where each image or news item is labeled by only some annotators, L2CU first estimates a consensus label per sample using Crowdlab, so no ground-truth training labels are needed. Each annotator is then converted into a fixed-length vector: for each class, L = 20 of that annotator's labels for samples of that class are sampled and concatenated in a consistent class order (annotators with fewer than 20 labels per class are dropped). These vectors are clustered with Fuzzy K-Means, with the number of clusters K chosen by silhouette score, and each annotator is assigned to the cluster with the highest membership. Each cluster becomes an "annotator profile."
Handling sparsity. Because each profile covers only a fraction of the data, L2CU estimates a C×C label transition matrix per profile capturing the probability that a profile member emits label n for a sample with consensus label c, and then samples G new noisy labels per sample from that matrix. This produces an augmented training set per profile that keeps the profile's error signature.
Cooperative model. Each profile gets its own model with three parts: a base model (any architecture, here a pretrained backbone) producing logits from the input; a human label encoder (a two-layer MLP with ReLU) embedding the user's label; and a decision model (a three-layer MLP) that concatenates both and outputs a class distribution. Training minimizes cross-entropy against consensus labels plus a weighted (λ) forward-correction term that maps clean predictions into the profile's noisy label space.
Test time. An unseen user labels a small clean validation set of M samples per class (M × C total, non-overlapping with training and test data), which is formatted the same way as the training annotator vectors and classified into a profile by an OVA SVM trained on training annotators. An entry condition compares the base model's validation accuracy against the user's; the cooperative model is used only if the base model is better. The paired model then runs on a held-out clean test set, and the positive/negative alteration metrics measure how its predictions changed the user's labels.
Experimental scale. Datasets are CIFAR-10 (50000 training, 200 validation, 9800 testing images, 10 classes; CIFAR-10N adds 747 annotators with three labels per image; CIFAR-10H adds 2571 annotators with about 51 labels per image on the test set), Fashion-MNIST-H (885 annotators, about 66 labels per image), Chaoyang (four-class pathology, 4021 training / 80 validation / 2059 testing images, three expert labels each in training) and AgNews for text. Simulated CIFAR-10 experiments flip 60% of samples in two classes across three profiles (airplane↔bird, horse↔deer, truck↔automobile), with five training and five testing users per profile. For CIFAR-10N, 159 annotators with at least 20 labels per class were split into 79 training and 80 testing users; Fashion-MNIST-H used 366 such annotators split 183/183. Training ran 500 epochs with early stopping after 20 unimproved epochs, λ = 0.1, ImageNet-1K pretrained backbones, Adam and NAdam optimizers, in PyTorch on an NVIDIA RTX 4090.
Why This Matters
Impact on research. The paper shifts the L2C conversation from "complement the users you trained on" to "complement users you have never met," which is the realistic deployment condition for any crowd-labeled or expert-labeled dataset. It also contributes a measurement tool — positive and negative alteration rates — that makes the trade-off between correcting people and corrupting their correct answers explicit, something plain accuracy hides. Its claim that a model can be correct when both the human and the AI are individually wrong is a notable, testable property.
Real-world applications:
- Medical diagnosis, where different pathologists have different characteristic confusions (the Chaoyang dataset, 0.858 → 0.988 in the reported results) and a model could correct errors specific to the reader it is paired with.
- Crowd-sourced image labeling for everyday objects, where thousands of annotators each have distinct biases (CIFAR-10N with 747 annotators).
- News and content categorization, where the paper demonstrates the approach transfers to text with AgNews (all 15 simulated test users improved).
- Fashion/retail product tagging, where large pools of annotators produce sparse noisy labels (Fashion-MNIST-H, 885 annotators, the largest reported gain at approximately 32%).
Industry relevance. Labeling platforms, data-quality teams, and any organization combining human review with automated classifiers could use the approach, since it needs no clean training labels and requires only M samples per class from a new user to adapt. The model-agnostic design means it can be layered on existing backbones — the reported results are stable across ResNet-50, DenseNet-121 and ViT/B-16.
Future Directions
- Improving profile matching in large annotator pools. Silhouette scores on the big real datasets are very low (0.01 on CIFAR-10N and CIFAR-10H), and the paper's own ablation shows that injecting SVM profiling errors increases not-improved users and lowers post-alteration accuracy. Better matching under weakly separated profiles is an open problem.
- Choosing K when silhouette scores are uninformative. The K ablation shows accuracy rising from K = 1 to the silhouette-optimal K = 2 on CIFAR-10N and dropping for K = 3, 6 and 10 due to over-adaptation as training users per profile shrink; more reliable model-selection criteria would help.
- Extending text-domain evaluation to real annotators. The AgNews result is a simulation; real multi-annotator text datasets with the sparsity properties L2CU requires are not evaluated in the reported content.
- Understanding the joint-error correction effect. The phenomenon that cooperation is correct when both human and AI are wrong is documented with a probability condition but its full characterization is left to the truncated discussion section, making it a clear avenue for further theoretical and empirical work.
Target Audience
Researchers and practitioners working on human-AI cooperation, learning-to-defer and learning-to-complement systems, and noisy-label or multi-annotator learning. It is also relevant to applied machine learning engineers building production systems that combine automated classifiers with human reviewers, and to dataset and annotation teams dealing with sparse, noisy labels from large pools of annotators. Readers without a background in cooperative classification will need the intermediate-level concepts of label transition matrices, clustering, and one-versus-all SVM classification.
Authors’ abstract
Recent research highlights the potential of machine learning models to learn to complement (L2C) human strengths; however, generalizing this capability to unseen users remains a significant challenge. Existing L2C methods oversimplify interaction between human and AI by relying on a single, global user model that neglects individual user variability, leading to suboptimal cooperative performance. Addressing this, we introduce L2CU, a novel L2C framework for human-AI cooperative classification with unseen users. Given sparse and noisy user annotations, L2CU identifies representative annotator profiles capturing distinct labeling patterns. By matching unseen users to these profiles, L2CU leverages profile-specific models to complement the user and achieve superior joint accuracy. We evaluate L2CU on datasets (CIFAR-10N, CIFAR-10H, Fashion-MNIST-H, Chaoyang and AgNews), demonstrating its effectiveness as a model-agnostic solution for improving human-AI cooperative classification.