Research
Prediction-Powered Semi-Supervised Learning with Online Power Tuning
Prediction-Powered Semi-Supervised Learning with Online Power Tuning Overview Research area: Semi-supervised learning (SSL), statistical inference, and online learning — specifically, extending the Pr

- arXiv
- 2510.22586
- Published
- 2025-10-26
- Authors
- Noa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy, Yaniv Romano
AI summary
Prediction-Powered Semi-Supervised Learning with Online Power TuningOverview
Research area: Semi-supervised learning (SSL), statistical inference, and online learning — specifically, extending the Prediction-Powered Inference (PPI) framework from parameter estimation to model training.
Technical level: Advanced. The paper combines a novel unbiased gradient estimator with finite-time non-convex convergence analysis, AdaGrad regret bounds, and an online tuning scheme for a one-dimensional interpolation parameter.
One-sentence scope: The paper proposes PP-SSL, an unbiased semi-supervised training framework whose interpolation weight λ between supervised and pseudo-label-corrected gradients is tuned on the fly with online learning rather than fixed offline.
Authors and venue: Noa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy, and Yaniv Romano (Department of Electrical and Computer Engineering, Technion IIT; Yaniv Romano also Department of Computer Science, Technion IIT). arXiv:2510.22586v1 [cs.LG], 26 Oct 2025. Code is available at https://github.com/noashoham/PP-SSL.
What This Paper Is About
Semi-supervised learning improves models by adding pseudo-labeled unlabeled data to a small set of trusted labeled data, but when the teacher model that generates those pseudo-labels is wrong — for example on an under-represented subgroup — the pseudo-labels inject bias into training and the model can end up worse than one trained on labeled data alone. Prediction-Powered Inference offers a debiasing correction, and PPI++ adds an interpolation parameter λ that balances the debiased pseudo-label term against the labeled-only term. The paper's goal is to make that correction usable during actual training: it derives an unbiased "prediction-powered" gradient and tunes λ dynamically during optimization so that the method matches the performance of the theoretically optimal λ, which depends on unknown quantities.
Key Contributions
-
A prediction-powered gradient estimator for SSL training. The authors build a gradient estimator g_PP^λ = g^n + λ(g̃^{N,f} − g^{n,f}) from labeled gradients, unlabeled pseudo-label gradients, and labeled pseudo-label gradients, and prove it is unbiased (E[g_PP^λ] = ∇L(w)) for any λ, since the labeled and unlabeled data share the same feature distribution.
-
Finite-time convergence guarantees rather than asymptotic ones. The paper analyzes the variance of the estimator, derives the variance-minimizing λ*, and shows that the reduced variance translates directly into a faster convergence rate of order O(√(V*/T) + 1/T) for smooth non-convex objectives. The authors contrast this with prior PPI-based SSL work (references [14, 27]) that provides only asymptotic guarantees at the population optimum, with no finite-time convergence bounds.
-
Online tuning of the interpolation parameter. Because λ* depends on the unknown labeled-gradient variance σ² and pseudo-label error variance σ_e², the authors treat λ as a trainable one-dimensional parameter updated with AdaGrad on h_t(λ) = ‖g_t + λ d_t^f‖², while the model weights w are also updated with an AdaGrad step size. Theorem 3.5 shows the online scheme incurs only an extra term of order √(Mβ)G/T, which decays faster than the leading O(√(V*/T)) term, so the overall rate matches that of the optimal λ*.
-
Empirical validation across synthetic and real data. Experiments on a synthetic linear regression problem, the California Housing tabular dataset, and the UTKFace facial age estimation dataset, covering both regression and classification metrics, show gains over classic SSL and over PPI-style training in settings where the teacher performs poorly on a subgroup.
Main Findings
-
Variance decreases with teacher quality and unlabeled data volume. Lemma 3.2 bounds n/4 · V(g_PP^λ) ≤ (1−λ)²σ² + λ²(rσ² + (1+r)σ_e²), where r = n/N is the ratio of labeled to unlabeled samples. The optimal value is λ* = (1/(1+r)) · σ²/(σ² + σ_e²).
-
The optimal variance bound behaves as expected in both regimes. With V* = (4σ²/n) · [((1 + σ²/σ_e²)^{-1} + r)/(1+r)]: when pseudo-labels are accurate (σ_e² ≪ σ²) the bound approaches (σ²/n)·(4r/(1+r)), a substantial reduction when r ≪ 1; when pseudo-labels are highly unreliable (σ_e² ≫ σ²) the bound approaches 4σ²/n, recovering standard labeled-only gradient variance up to a factor of 4, which the authors describe as asymptotically equivalent.
-
Unreliable pseudo-labels are down-weighted automatically rather than biasing the model. The authors emphasize that this contrasts with conventional pseudo-labeling, which may reduce empirical variance but typically introduces bias when pseudo-labels are incorrect.
-
Pseudo-label error variance is controlled by teacher prediction error. Lemma 3.1 shows that if ∇ℓ(w;x,y) is L_Y-Lipschitz in y, then σ_e² ≤ L_Y² · E^f, where E^f = E(y − f(x))² is the teacher's prediction error. The Lipschitz condition is shown in Appendix A to apply to squared and logistic losses, among others.
-
Online tuning matches the oracle λ. The convergence bound in Theorem 3.5 is O(√(MβV*/T) + Mβ/T + √(Mβ)G/T) with η_0 = √(2M/β); the third, online-tuning term decays faster than the first, so the method asymptotically matches the rate achievable with the optimal but unknown λ*.
-
Synthetic results: PP-SSL beats SSL and PPI++ under teacher bias. Across μ ∈ {0.1, 1, 3, 5, 7} (bias magnitude) with n = 20, N = 1,000, and n_test = 1,000, evaluated over 100 independent experiments, average MSE rises with μ, but PP-SSL performs best. In group B, where the teacher is poor, the SSL model does worse than the Only Labeled model, while PP-SSL tends to achieve the best MSE for that group, especially in high-bias regimes. Group A is where the teacher is oracle-like by design, so SSL does well there.
-
Fixed-λ comparison. Figure 2 compares PP-SSL's final test MSE against a PPI++-inspired baseline using constant λ. The adaptive method achieves performance comparable to the baseline with the best fixed λ, and the preferred λ shrinks as teacher bias grows, consistent with the theory.
-
California Housing results. With n ≈ 100, N ≈ 18,000, n_val ≈ 1,000, n_test ≈ 1,000, a teacher trained on N_A ≈ 50 group A samples and N_B in the range 5 to about 50 group B samples, and 100 data splits, MSE falls for all methods as N_B/N_A grows and the teacher improves on group B. PPI++ and PP-SSL have comparable MSE overall, both tending to outperform the other baselines when the teacher becomes less accurate due to limited group B data.
-
UTKFace results. Using n ≈ 700, N ≈ 16,500, n_val ≈ 2,000, n_test ≈ 2,000, a ResNet50, a teacher trained on N_A ≈ 800 samples from group A (ages over 30) and N_B from group B (ages 0-29) ranging from 10 to about 800, the SSL method proves highly sensitive to teacher quality — performing worse than Only Labeled — while PPI++ and PP-SSL achieve lower MSE, with the proposed method showing a noticeable advantage.
-
Truncation note. The supplied paper content ends mid-sentence in Section 4.3.1; the results of Section 4.3.2, the classification-task results, the full numerical tables, and the appendix experiments (including Appendix C's comparison of λ* with PPI++'s λ for linear regression, Appendix D's no-group-indicator setting, Appendix E's additional California Housing experiments, and Appendix H's analysis of λ dynamics) are not reported in the available text.
Methodology in Plain English
The authors set up a teacher-student training loop. In each round the learner receives a small batch of n labeled examples, a large batch of N unlabeled examples (with N ≫ n), and a fixed teacher model f that produces pseudo-labels. From these, three gradients are computed: the ordinary labeled gradient, the unlabeled pseudo-label gradient, and the labeled pseudo-label gradient. The last two are combined as a correction term scaled by λ and added to the labeled gradient, giving an estimate that is unbiased for the true objective no matter how bad the teacher is — because the pseudo-label loss on labeled data has the same expectation as the pseudo-label loss on unlabeled data, cancelling the teacher's error.
The first part of the analysis asks: what value of λ minimizes the estimator's variance? The answer depends on σ² (noise in labeled gradients), σ_e² (noise in the loss-error gradients), and the ratio r = n/N, which are all unknown in practice. The second part solves this by making λ a second trainable quantity. The authors use AdaGrad, whose regret bound depends on the sum of squared gradient norms, updating λ on the convex function h_t(λ) = ‖g_t + λ d_t^f‖² so that the cumulative second moment — equivalent to cumulative variance, since the estimator is conditionally unbiased — is minimized. The model weights use an AdaGrad step size too, which lets the algorithm adapt to σ² and σ_e² implicitly without knowing them. The resulting algorithm is one extra scalar update per training iteration.
Empirically, the authors construct or select settings where the population is split into a subgroup the teacher handles well and a subgroup it handles badly: a binary group indicator in the synthetic data, the 40% lowest house prices in California Housing, and ages 0-29 in UTKFace. They compare against the teacher itself, a labeled-only model, a standard pseudo-labeling SSL model, and a PPI++ model with the offline λ from prior work, reporting MSE, MAE, R² for regression and accuracy for classification.
Why This Matters
Impact on research. The paper moves PPI-style debiasing from the offline inference setting into iterative training with finite-time guarantees. It shows that the variance argument behind PPI++ survives contact with the actual optimization trajectory, not just the population optimum that is never reached during training. It also reframes "how much should I trust my pseudo-labels?" as an online convex optimization problem, which opens a route to adaptively weighting other sources of noisy supervision (weak labels, noisy annotators, teacher ensembles) during training rather than beforehand.
Real-world applications:
- Clinical and biomedical modeling, where expert labels are scarce but unlabeled records are abundant, and where a pretrained model may be systematically worse for an under-represented patient subpopulation.
- Demographic subgroup fairness in tabular prediction, such as housing or credit scoring, where a minority subgroup with few labels would otherwise be overwhelmed by abundant but biased pseudo-labels.
- Visual regression on long-tailed attribute distributions, as demonstrated with facial age estimation over ages 0-29 and over 30, where a teacher trained mostly on adults mislabels young faces.
- Weak-supervision pipelines more generally, where labelers or heuristics are cheap and plentiful but unevenly accurate across regions of the input space.
Industry relevance. The method requires no extra data collection, no knowledge of teacher accuracy, and no held-out tuning of a key hyperparameter — only one additional scalar variable updated with a standard AdaGrad rule. That makes it a low-overhead drop-in correction for existing teacher-student training pipelines where the alternative is a costly hyperparameter sweep whose optimal value depends on unknown quantities.
Future Directions
- Extending beyond the fixed-teacher assumption. The analysis relies on a teacher whose predictions remain constant while the student trains. Whether the guarantees carry over to self-training, where pseudo-labels are refreshed from the student, is left open.
- Applying the online-tuning idea to other interpolation and weighting parameters in semi-supervised objectives, since the framework only requires the per-round loss h_t(λ) to be convex in the tuned parameter.
- Characterizing behavior beyond the bounded-gradient and bounded-objective assumptions used in Theorem 3.5, which the paper states as requirements for the guarantee.
- Releasing the full experimental picture. The available content stops partway through the visual experiments and does not include the second real visual dataset, the classification results, the comparison with PPI++'s λ in the linear-regression setting (Appendix C), the no-group-indicator setting (Appendix D), or the λ dynamics study (Appendix H).
Target Audience
This paper is most useful for machine learning researchers working on semi-supervised learning, weak supervision, or statistical inference with pseudo-labels, and for theoretically inclined graduate students comfortable with smooth non-convex convergence rates, AdaGrad regret bounds, and bias-variance decompositions of gradient estimators. Practitioners who train teacher-student pipelines on data with scarce labels and uneven teacher quality will find the algorithm itself compact and implementable, but they will need to work through the variance analysis to appreciate why the online λ update is preferable to a fixed value. Readers looking for a broad empirical benchmarking study will find the paper's experiments targeted at the specific failure mode — a subgroup where the teacher is inaccurate — rather than at general SSL leaderboards.
Authors’ abstract
Prediction-Powered Inference (PPI) is a recently proposed statistical inference technique for parameter estimation that leverages pseudo-labels on both labeled and unlabeled data to construct an unbiased, low-variance estimator. In this work, we extend its core idea to semi-supervised learning (SSL) for model training, introducing a novel unbiased gradient estimator. This extension addresses a key challenge in SSL: while unlabeled data can improve model performance, its benefit heavily depends on the quality of pseudo-labels. Inaccurate pseudo-labels can introduce bias, leading to suboptimal models.To balance the contributions of labeled and pseudo-labeled data, we utilize an interpolation parameter and tune it on the fly, alongside the model parameters, using a one-dimensional online learning algorithm. We verify the practical advantage of our approach through experiments on both synthetic and real datasets, demonstrating improved performance over classic SSL baselines and PPI methods that tune the interpolation parameter offline.