Skip to content
AI.info

Research

Learning from Interval Targets

Learning from Interval Targets — Paper Summary Overview Research area: Machine learning theory and methodology, specifically regression under weak supervision where labels are given as intervals rathe

Learning from Interval Targets
arXiv
2510.20925
Published
2025-10-23
Authors
Rattana Pukdee, Ziqi Ke, Chirag Gupta

AI summary

Learning from Interval Targets — Paper Summary

Overview

Research area: Machine learning theory and methodology, specifically regression under weak supervision where labels are given as intervals rather than exact values.

Technical level: Advanced. The paper is primarily a statistical learning theory contribution, relying on Rademacher complexity, Lipschitz continuity of hypothesis classes, non-asymptotic generalization bounds, and a minmax (adversarial worst-case) formulation. Readers should be comfortable with empirical risk minimization and generalization theory.

One-sentence scope: The paper develops and analyzes two methods — a projection loss and a minmax loss — for training regression models when each training target is only known to lie within an interval ([l, u]).

What This Paper Is About

In many real-world settings, the exact target value for a regression problem cannot be observed; only a lower and upper bound on it is available (for example, due to expensive medical measurements, sensors that record values only at discrete intervals, or human labelers who can only give ranges). This paper asks how to train a regression model from such interval targets, and what theoretical guarantees can be established without the restrictive assumptions (realizability and a "small ambiguity degree") that prior work relied on.

Key Contributions

  1. A generalization bound for the projection loss that removes the ambiguity-degree assumption. The authors show that for a hypothesis class with Rademacher complexity decaying as (O(1/\sqrt{n})) (such as two-layer neural networks with bounded weights), the error decomposes into an irreducible term depending on interval quality and the Lipschitz constant of the hypothesis class, plus terms that vanish at (O(1/\sqrt{n})). This is non-asymptotic, applies in the agnostic setting, and reveals how hypothesis class structure affects learning.

  2. A key structural insight about smooth hypotheses shrinking intervals. When the hypothesis class is smooth (m-Lipschitz), outputs for nearby inputs cannot differ much, which rules out portions of the original intervals and yields much smaller valid intervals (Proposition 3.4 for the zero-projection-loss case, and Theorem 3.6 for the approximate case).

  3. A minmax learning formulation with two variants. The paper minimizes loss against the worst-case target within each interval. Variant (i) allows worst-case labels to be any point in the interval; variant (ii) restricts them to outputs of some hypothesis in the class, incorporating smoothness. Proposition 5.4 shows there are scenarios where the second variant performs arbitrarily better — with (\operatorname{err}(f_1)=0) while (\operatorname{err}(f_2)>c) for any constant (c>0).

  4. Experiments on real-world datasets. The authors report that both methods achieve state-of-the-art performance, though the truncated content does not name the datasets or report specific numbers.

Main Findings

  • The projection loss is a valid and efficiently computable proxy. For any loss satisfying (\ell(y,y')=\psi(|y-y'|)) with (\psi) non-decreasing and zero only at equality, the projection loss (\pi_\ell(f(x),l,u)=\min_{\tilde{y}\in[l,u]}\ell(f(x),\tilde{y})) equals zero if and only if (f(x)\in[l,u]), and can be written using only the interval boundaries: (1[f(x)<l]\ell(f(x),l)+1[f(x)>u]\ell(f(x),u)) (Proposition 2.1).

  • Smoothness shrinks intervals for perfectly-fitting hypotheses. For any (f) in (\widetilde{\mathcal{F}}0) (expected projection loss zero) within an m-Lipschitz class, (f(x)) lies in the intersection of all intervals induced by other points, ([l{\mathcal{D}\to x}^{(m)}, u_{\mathcal{D}\to x}^{(m)}]) (Proposition 3.4). This intersection is always smaller than the original ([l_x,u_x]) and shrinks further as smoothness increases — effectively denoising the intervals.

  • Buffer terms quantify how much smoothness alone can be relaxed. For (f\in\widetilde{\mathcal{F}}\eta) (expected projection loss at most (\eta)) with (\ell(y,y')=|y-y'|^p), (p\ge1), the reduced interval must be expanded by buffers (r\eta(x)) and (s_\eta(x)) defined implicitly by expectations over the lower and upper bound gaps (Theorem 3.6). Proposition 3.7 bounds these buffers in terms of (\eta) and the probability that the bound gap is small.

  • The realizability bound splits into irreducible and vanishing parts. In the realizable setting (Theorem 4.1), (\operatorname{err}(f)) is bounded by an irreducible term (\mathbb{E}X[|u{\mathcal{D}\to X}^{(m)}-l_{\mathcal{D}\to X}^{(m)}|]) that depends on interval quality and smoothness and does not decrease with (n), plus terms (\tau + (D/\sqrt{n}+M\sqrt{\ln(1/\delta)/n}),\Gamma(\tau)) that decay to zero as (n\to\infty) for any fixed (\tau). The function (\Gamma(\tau)) is decreasing in (\tau) and depends on the distribution of intervals (\mathcal{D}_I).

  • The agnostic bound retains a floor. In the agnostic setting (Theorem 4.2), the bound adds (\operatorname{OPT}) (the error of the best hypothesis in (\mathcal{F})) and does not converge to zero as (n\to\infty); it converges to (\operatorname{OPT} + \mathbb{E}X[|u{\mathcal{D}\to X}^{(m)}-l_{\mathcal{D}\to X}^{(m)}|] + 2\tau + 2,\operatorname{OPT}\cdot\Gamma(\tau)). The optimal (\tau) satisfies (\tau = \operatorname{OPT}\cdot\Gamma(\tau)).

  • Trade-off in choosing smoothness. A smaller Lipschitz constant (m) reduces the irreducible interval term, but if (m) is too small the class may not contain a good hypothesis, making (\operatorname{OPT}) large. The authors suggest tuning (m) as a hyperparameter on a validation set.

  • The mid-point heuristic is a special case of the minmax loss. For (\ell(y,y')=|y-y'|), (\rho_\ell(f(x),l,u)=|f(x)-\frac{l+u}{2}|+\frac{u-l}{2}), so minimizing the worst-case loss is equivalent to supervised learning with the interval midpoint as the target (Corollary 5.2).

  • Constraining worst-case labels to the hypothesis class is provably better. The minmax objective against (\widetilde{\mathcal{F}}0) is a tighter upper bound than the general worst-case (\rho\ell) bound (Proposition 5.3), and Proposition 5.4 shows an arbitrary gap in performance between the two variants in the worst case.

  • Prior work's assumptions are limited. The realizability assumption requires (f^*\in\mathcal{F}); the small ambiguity degree assumption requires the intersection of infinitely many sampled intervals to be the singleton ({y}), which fails even for the simple case ([l,u]=[y-\epsilon,y+\epsilon]) since (y+\epsilon/2) also always lies in the interval. The authors argue that ambiguity degree is less suitable for regression, where predictions merely close to the target (within tolerance (\epsilon)) are often acceptable, and explore an "ambiguity radius" extension in the appendix (Section F).

  • What the paper does not report. The truncated content does not name the real-world datasets, benchmark names, dataset sizes, or any numeric experimental results.

Methodology in Plain English

The authors begin from a simple idea: since the true label is guaranteed to lie inside ([l,u]), train a model whose output falls inside the interval. The 0-1 version of this loss is discontinuous, so they use a projection loss that penalizes the output only when it falls outside the interval, penalizing by the distance to the nearer boundary. This loss is equivalent to the partial-label learning loss used in classification (Lv et al., 2020) and generalizes the limiting method of Cheng et al. (2023a).

Theoretically, they first characterize the set (\widetilde{\mathcal{F}}\eta) of hypotheses with expected projection loss at most (\eta). For (\eta=0), every point's output must be inside its interval; combining this with an m-Lipschitz constraint, each point's interval can be narrowed using information from every other point's interval (since a nearby point's interval shifts by at most (m|x-x'|)). This gives the reduced interval of Proposition 3.4. For (\eta>0), the output may lie outside its interval, so they introduce "bound gaps" measuring how much each point's induced bounds differ from the tightest bounds, and add compensating buffers (r\eta(x), s_\eta(x)) whose size is determined by how the projection loss budget (\eta) interacts with those gaps. Plugging these reduced intervals into a standard uniform-convergence argument (Rademacher complexity) yields the generalization bounds of Section 4.

Separately, they consider a minmax loss: instead of the best point in the interval, penalize the worst point. For general losses this equals evaluating the loss at the upper bound when the prediction is below the interval midpoint, and at the lower bound otherwise. They then propose a stronger variant where the adversary is restricted to hypotheses in (\widetilde{\mathcal{F}}_0), i.e., functions that themselves fit all intervals, so the adversary's choices respect smoothness rather than being arbitrary points. An empirical version of this objective is given at the end of Section 5.

Why This Matters

Impact on research. This work moves interval-target regression from asymptotic guarantees under strong assumptions (realizability plus small ambiguity degree) to non-asymptotic guarantees in the agnostic setting, and shows explicitly how the smoothness of the hypothesis class interacts with the quality of the intervals. It also connects the practical "use the midpoint" heuristic to a principled worst-case objective, and offers a provably tighter alternative based on smoothness.

Real-world applications (as described in the paper):

  • Medical measurements where exact values are expensive or invasive to obtain.
  • Sensor deployments that record a target only at discrete intervals (e.g., every hour), leaving intermediate values unobserved.
  • Human labeling tasks where annotators can more easily provide a range than a precise value — a form of weak supervision.
  • Bond pricing, where domain knowledge or data properties make intervals readily available as side information.

Industry relevance. The Bloomberg affiliations and the internship note suggest direct applicability to financial modeling, where ranges and uncertainty bounds are common. More broadly, the framework applies to any regression pipeline where label precision is limited by cost, instrument resolution, or privacy.

Future Directions

  1. Choosing the smoothness level in practice. The theory shows a trade-off between a smaller Lipschitz constant (tighter interval term) and the risk of a worse best-in-class hypothesis ((\operatorname{OPT})); the paper suggests tuning (m) on a validation set but does not report a systematic procedure.

  2. The ambiguity-radius extension for regression. Section 2.4 of the related work defers an extension of "ambiguity degree" to an "ambiguity radius" to Section F, arguing the paper's analysis is stronger there — a direction the authors flag as an open comparison.

  3. Probabilistic interval setting. The main paper focuses on deterministic intervals (one fixed ([l_x,u_x]) per (x)); the general probabilistic setting where multiple intervals are drawn per input is deferred to Appendix D and remains less fully developed in the main text.

  4. Efficiently solving the smoothness-constrained minmax objective. The empirical objective restricted to worst-case labels from (\widetilde{\mathcal{F}}_0) is stated but the truncated content does not describe optimization algorithms or guarantees for it, leaving a gap between the theory and a practical solver.

Target Audience

This paper is best suited for machine learning researchers and graduate students working on weak supervision, partial-label learning, learning theory, and generalization bounds, as well as practitioners in domains with costly or imprecise labels (healthcare, finance, sensor-based systems) who need principled methods and want to understand when interval-supervised regression can be trusted. A working knowledge of statistical learning theory is required to follow the main results; the experimental section, if read alone, is more accessible.

Authors’ abstract

We study the problem of regression with interval targets, where only upper and lower bounds on target values are available in the form of intervals. This problem arises when the exact target label is expensive or impossible to obtain, due to inherent uncertainties. In the absence of exact targets, traditional regression loss functions cannot be used. First, we study the methodology of using a loss functions compatible with interval targets, for which we establish non-asymptotic generalization bounds based on smoothness of the hypothesis class that significantly relaxing prior assumptions of realizability and small ambiguity degree. Second, we propose a novel min-max learning formulation: minimize against the worst-case (maximized) target labels within the provided intervals. The maximization problem in the latter is non-convex, but we show that good performance can be achieved with the incorporation of smoothness constraints. Finally, we perform extensive experiments on real-world datasets and show that our methods achieve state-of-the-art performance.

Read the original paper