Research
Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics
Summary: Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics Overview Research area: Statistical machine learning and nonparametric statistics, specifically kernel-based tw
- arXiv
- 2512.13997
- Published
- 2025-12-16
- Authors
- Aaron Wei, Milad Jalali, Danica J. Sutherland
AI summary
Summary: Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-StatisticsOverview
Research area: Statistical machine learning and nonparametric statistics, specifically kernel-based two-sample testing with Maximum Mean Discrepancy (MMD), generalized U-statistics theory, and kernel selection for test power optimization.
Technical level: Advanced. The paper relies on reproducing kernel Hilbert spaces, Hilbert–Schmidt operators, covariance operators, mean embeddings, and asymptotic theory of (generalized) U-statistics, including degenerate limiting distributions expressed as infinite mixtures of shifted chi-squared variates.
One-sentence scope: The paper extends generalized U-statistics theory to characterize the asymptotic distribution and variance of the standard MMD estimator when the two samples have unequal sizes — including non-proportional regimes — and uses this to derive a criterion for choosing kernels that maximize test power without discarding data.
Authors and affiliations: Aaron Wei, Milad Jalali, and Danica J. Sutherland (University of British Columbia; Independent researcher, Vancouver, Canada; University of British Columbia & Alberta Machine Intelligence Institute). Reviewed on OpenReview: https://openreview.net/forum?id=KjXW75GHHF. License: CC BY 4.0.
What This Paper Is About
Two-sample testing asks whether two distributions P and Q differ, using only samples drawn from each. MMD-based tests are popular and powerful, but the theory used to choose a good kernel (and hence maximize test power) has largely assumed both samples are the same size. When sample sizes are unequal, the usual MMD estimator is no longer an ordinary U-statistic, so the established asymptotic machinery does not directly apply — and practitioners have often responded by throwing away data to force equal sizes.
This paper closes that gap by developing the asymptotics of the MMD estimator through the lens of generalized U-statistics, covering both proportional and non-proportional sample-size regimes, and showing that the resulting procedure reduces to the standard approach when sample sizes are equal.
Key Contributions
-
A generalization of generalized U-statistics asymptotics. The paper fills in results for the asymptotic behavior of generalized U-statistics — including cases with more than two samples (c > 2) and arbitrary block sizes m_j — building on and substantially extending classical theorems.
-
Asymptotic distributions of the MMD estimator under unequal sample sizes. The authors derive the limiting distribution of the standard unbiased MMD estimator under the null hypothesis (MMD²(P,Q) = 0) and under the alternative, scaling by min(n_X, n_Y) rather than n_X + n_Y. This permits regimes such as n_Y = n_X² and n_Y = 5√n_X, which previous proportional-regime results excluded.
-
Cleaner variance characterizations and a corrected degeneracy picture. The paper gives simplified closed forms for the variance of MMD estimators, and shows that the estimator can be degenerate (first-order degenerate) even when the MMD is nonzero — contradicting informal claims in several prior papers. It constructs such a case and proves it cannot occur in common settings.
-
A new power-optimization criterion for unequal sample sizes. The derived distributions yield a finite-sample estimate of the signal-to-noise ratio that can be maximized to select a kernel, replacing the wasteful practice of training on equal-sized subsamples.
Main Findings
-
min(n_X, n_Y) is the correct scaling, not n_X + n_Y. Experiments with P = Q = Laplace(0, 1/√2) and a unit-lengthscale Gaussian kernel show that in the proportional case (n_Y = n_X/2) both scalings converge, since n_X + n_X/2 = 3·min(n_X, n_X/2). In the non-proportional case (n_Y = 5√n_X), the n_X + n_Y scaling clearly fails to converge in distribution while the min(n_X, n_Y) scaling does, matching the limiting distribution in the paper's null-case theorem.
-
The alternative-hypothesis case confirms the same scaling. With P = Laplace(0, 1/√2) and Q = Laplace(0, 3) under a unit-lengthscale Gaussian kernel, and MMD(P,Q) > 0 estimated from 10000 samples, the min(n_X, n_Y) scaling converges as predicted while n_X + n_Y does not in the non-proportional setting. Quantiles of the limiting distribution were computed from a variance estimated on a sample with n_X = n_Y = 10000.
-
Degeneracy is not equivalent to zero MMD. The paper proves that the U-statistic MMD estimator is infinitely degenerate (zero variance) if and only if both covariance operators C_P and C_Q are zero — which can happen even for distinct distributions, such as distinct point masses. First-order degeneracy occurs whenever μ_P = μ_Q. When μ_P ≠ μ_Q, the estimator may still be first-order degenerate, contrary to claims in prior work.
-
Common settings rule out the surprising case. If the supports of P and Q are not disjoint, with continuity of k(x, ·) and finite sup_x k(x, x), the estimator is degenerate if and only if μ_P = μ_Q. The same "if and only if" holds under real-analytic kernels on ℝ^d with finite sup_x k(x, x) and supports of positive Lebesgue measure; this remains true if only one of P, Q has a support of positive Lebesgue measure. Under these conditions, the positivity of ⟨μ_P − μ_Q, C_P(μ_P − μ_Q)⟩ is equivalent to that of ⟨μ_P − μ_Q, C_Q(μ_P − μ_Q)⟩.
-
Closed-form variance with equal sample sizes. For n_X = n_Y = n, the variance of the U-statistic MMD estimator simplifies to (4/n)⟨μ_P − μ_Q, (C_P + C_Q)(μ_P − μ_Q)⟩_H + (2/(n(n−1)))·‖C_P + C_Q‖²_HS.
-
Closed-form variance with unequal sample sizes. The paper gives an explicit expression for the variance of MMD² for any n_X, n_Y ≥ 2, composed of terms scaling as 1/n_X, 1/n_Y, 1/(n_X(n_X−1)), 1/(n_Y(n_Y−1)), 1/(n_X n_Y), and a correction involving the factor ((n_X−2)/(n_X−1))·((n_Y−2)/(n_Y−1)) multiplied by [⟨μ_P, C_P μ_P⟩ + ⟨μ_Q, C_Q μ_Q⟩]. Each large-parenthesis factor tends to 1 as n_X, n_Y → ∞ regardless of their relative rates.
-
The rate bound can be loose outside the proportional regime. For a generalized U-statistic of finite degeneracy order r, the variance is O(n_min^{−(r+1)}), which is tight (Θ) in the proportional regime n_i = Θ(n_min). In general only O holds: if C_P = 0, ⟨μ_P − μ_Q, C_Q(μ_P − μ_Q)⟩ > 0, and ⟨μ_Q, C_Q μ_Q⟩ > 0, the estimator is non-degenerate but its variance is Θ(1/n_Y) — which for n_Y = n_X¹⁰ is Θ(n_min^{−10}), not Θ(n_min^{−1}).
-
The null limiting distribution is an infinite mixture of centered chi-squared variates. Under MMD²(P,Q) = 0, with min{n_X, n_Y}/n_X → ρ_X and min{n_X, n_Y}/n_Y → ρ_Y in [0,1], min{n_X, n_Y}·MMD² converges in distribution to (ρ_X + ρ_Y) Σ_l λ_l (Z_l² − 1), where the Z_l are independent standard normals and the λ_l solve an integral equation involving the centered kernel φ(X) − μ_P. This extends Gretton et al. (2012, Theorem 12), which required n_X/(n_X + n_Y) → ℓ_X and n_Y/(n_X + n_Y) → ℓ_Y in (0,1) and scaled by n_X + n_Y.
-
Consistency and power. As n → ∞, any test with asymptotic level control is consistent whenever μ_P ≠ μ_Q, regardless of degeneracy order. For large samples, asymptotic power is dominated by the signal-to-noise ratio MMD²(P,Q)/σ_1(P,Q), motivating kernel selection by maximizing a finite-sample estimate of that quantity on a training set.
-
The estimator is nearly interchangeable with the U-statistic form. The two equal-size estimators differ only by omitted k(x_i, y_i) terms; with n samples each, |MMD² − MMD_U²| ≤ (8 sup_x k(x,x)/n^{3/2})·√(log(2/δ)) with probability at least 1 − δ.
-
MMD² has the same order of degeneracy as the U-statistic estimator, so the degeneracy characterization carries over.
Methodology in Plain English
The authors start from a known observation (also made by Kim et al. (2022) and Schrab et al. (2023)): the standard unbiased MMD estimator is a generalized U-statistic — an average over pairs drawn from each of two samples — with c = 2 groups and two arguments per group. Ordinary U-statistics theory assumes a single sample, so the authors work out the missing asymptotics for the generalized case.
They build on the classical variance decomposition for generalized U-statistics, which expresses the variance as a sum of conditional-variance terms indexed by how many arguments are conditioned on. Terms where no single argument carries information are degenerate, and the slowest-decaying non-zero term determines the estimator's rate and limiting distribution. They apply this to the MMD kernel specifically, simplify the resulting expression, and identify precisely when degeneracy occurs.
For the null distribution, they extend the classic approach of Serfling and Anderson et al. (later used for MMD in the proportional setting by Gretton et al. (2012)) to the generalized, unequal-size, non-proportional setting, using an orthogonal expansion to handle asymmetric kernels. Since the general result is complicated, they present the MMD-specific case explicitly.
They verify the theory empirically with simulation studies: histograms and inset Q-Q plots comparing empirical quantiles of the scaled estimator to the predicted limiting distribution. Eigenvalues were estimated from a sample with n_X = n_Y = 5000, and the alternative-case quantiles from a variance estimated with n_X = n_Y = 10000.
Finally, they show the new test reduces to the previous equal-sample-size framework, and derive from the asymptotic power expression a kernel-selection criterion that uses all available data.
Why This Matters
Impact on research. Kernel selection for MMD tests has relied on asymptotics valid only for equal (or proportionally sized) samples, which pushed practitioners toward discarding data. This work removes that restriction, provides clean variance formulas that replace previously cumbersome derivations, and corrects an error repeated (informally) in the literature about when the estimator is degenerate. The generalized U-statistics results are of independent interest beyond MMD.
Real-world applications cited in the paper:
- Distinguishing treatment and control groups (e.g., Kobayashi et al., 2017), where group sizes are rarely balanced.
- Evaluating and training generative models (e.g., Li et al., 2015; Dziugaite et al., 2015; Bińkowski et al., 2018; Jayasumana et al., 2024), where generated and reference sets often differ in size.
- Domain adaptation (e.g., Long et al., 2013), where labeled source and target samples are typically imbalanced.
- General scientific and industrial settings where data collection is expensive in one group and abundant in another, so discarding samples is costly.
Industry relevance. Practitioners of MMD-based testing in any setting with unbalanced data can keep all samples instead of subsampling to equal size. The paper also notes that permutation testing — not the asymptotic distribution — is typically how thresholds are set in practice, so the cost of kernel learning dominates computationally; that means the new power-optimization criterion plugs directly into existing kernel-learning workflows (including deep kernel selection), and could be used to learn kernels for a cross-MMD test if desired.
Future Directions
The truncated content does not include an explicit future-work section, so the following are open questions raised by the work rather than stated plans:
-
Extending the null-distribution theorem to the first-order degenerate case with nonzero MMD. The paper explicitly states that it does not handle first-order degeneracy when μ_P ≠ μ_Q, because it makes the proof more difficult; characterizing that distribution remains open.
-
Characterizing exactly when the loose O rate is not tight outside proportional regimes. The paper gives one example (C_P = 0 with n_Y = n_X¹⁰) where the bound is far from Θ; a full characterization of which (C_P, C_Q, μ_P, μ_Q, n_X, n_Y) combinations produce Θ versus only O would sharpen the theory.
-
Validating the new power criterion against competing kernel-selection methods. The paper contrasts its finite-sample signal-to-noise estimate with the wasteful subsampling approach of Liu et al. (2020) and notes that classifier-based tests (Lopez-Paz & Oquab, 2017; Kim et al., 2021; Kübler et al., 2022) are easier to use but often weaker; head-to-head power comparisons on realistic unbalanced benchmarks would test the proposed approach in practice.
-
Applying the kernel-learning criterion to cross-MMD tests. The paper observes that cross-MMD power is closely related to full-MMD power, so the new criterion could train kernels for cross-MMD, which avoids permutation testing — but the authors see little practical reason to do so, leaving the question of when it would be worthwhile unanswered.
Target Audience
This paper is aimed at researchers and graduate students in statistical machine learning, kernel methods, and nonparametric statistics who work on two-sample testing, kernel selection, or U-statistic asymptotics. It is also relevant to methodologists who need the technical justification for using MMD tests on unbalanced data, and to practitioners building deep-kernel or MMD-based evaluation pipelines for generative models, domain adaptation, or A/B-style comparisons where sample sizes are unequal.
Authors’ abstract
Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require discarding valuable data, unnecessarily reducing test power. We address this long-standing limitation by extending the theory of generalized U-statistics and applying it to the usual MMD estimator, resulting in new characterization of the asymptotic distributions of the MMD estimator with unequal sample sizes (particularly outside the proportional regimes required by previous partial results). This generalization also provides a new criterion for optimizing the power of an MMD test with unequal sample sizes. Our approach preserves all available data, enhancing test accuracy and applicability in realistic settings. Along the way, we give much cleaner characterizations of the variance of MMD estimators, revealing something that might be surprising to those in the area: while zero MMD implies a degenerate estimator, it is sometimes possible to have a degenerate estimator with nonzero MMD as well; we give a construction and a proof that it does not happen in common situations.