Research
Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
Overview Research area: Quantitative finance / portfolio theory and statistical text analysis (q-fin.ST), combining optimal transport theory, factor-risk models, and language-model text embeddings. Te

- arXiv
- 2608.29692
- Published
- 2026-08-30
- Authors
- Marcus Gawronsky, Chun-Sung Huang
AI summary
Overview
Research area: Quantitative finance / portfolio theory and statistical text analysis (q-fin.ST), combining optimal transport theory, factor-risk models, and language-model text embeddings.
Technical level: Advanced. The paper is a mathematical theory contribution with measure-theoretic optimal transport, Hilbert-space identities, and formal theorems, followed by a compact empirical exercise.
Scope: The paper derives one-sided, certifiable upper bounds on systematic portfolio variance from firm-level distribution-valued text characteristics, without estimating cross-asset return covariances, and turns that bound into a long-only allocation rule.
What This Paper Is About
Mean–variance portfolio construction requires a cross-asset return covariance matrix, but that matrix is hardest to estimate in short, high-dimensional panels: an unrestricted covariance for n assets has n(n+1)/2 entries, while a demeaned return history of length T has rank at most min(T−1, n). The authors ask a different question: how much portfolio risk can be ruled out using only observable firm-level information, before any covariances are estimated? They represent each firm by the distribution of its article embeddings, measure separation between firms with the quadratic Wasserstein distance, and convert that observed "information geometry" into a certified reduction in a perfect-positive-dependence variance benchmark — then minimize the resulting bound to obtain an implementable portfolio.
Key Contributions
-
A coherent, sharp portfolio-level variance certificate. The paper derives a multi-firm upper bound on systematic portfolio variance that keeps all asset exposure marginals under one joint law (Theorem 1, Theorem 3), rather than stacking incompatible pairwise couplings.
-
A computationally tractable pairwise relaxation used as the decision rule. A weighted pairwise construction, C(q) = ½ Σ_i Σ_j q_i q_j ℓ_ij², requires only marginal volatility scales and observed pairwise Wasserstein distances — no cross-asset return covariances — and is quadratic in the portfolio weights.
-
A carrier-invariance result. When firm-specific slack is zero, the common carrier scale L changes the certified variance reduction but not the normalized allocation, which depends only on the observed information geometry (Corollary 1).
-
An in-sample empirical exercise on a 52-firm panel, 2018–2022. An allocation built from Qwen3-Embedding-8B news representations is compared with equal risk weighting across four prespecified capped long-only portfolio populations.
Main Findings
-
Headline bound: For a standardized long-only portfolio with normalized risk weights q, the systematic variance obeys Var(R_q) ≤ 1 − C(q). The value 1 is the variance benchmark under perfect positive dependence, and C(q) is the portion the observed information geometry certifies away. The bound is a certificate about admissible joint risk configurations, not a point estimate of the covariance matrix.
-
Pairwise observable floors: The floor for assets i and j is ℓ_ij = [L⁻¹ W₂(C_i, C_j) − τ_i − τ_j]₊, where W₂(C_i, C_j) is the observed Wasserstein-2 separation of the firms' article-embedding distributions, L is the common antilipschitz carrier constant, and τ_i, τ_j are firm-specific slack radii. A zero floor means the maintained restrictions certify no positive separation for that pair.
-
Constant-capital-weight form: For capital weights x with marginal volatility scales σ_i, with A(x) = Σ_i x_i σ_i and q_i(x) = x_i σ_i / A(x), the raw-return certificate is Var(R_x) ≤ A(x)² {1 − C(q(x))}. A(x)² is the perfect-dependence variance benchmark; the residual sensitivity term δ is described as a portfolio residual-covariance budget for sensitivity analysis, not an estimated structural parameter. At δ = 0, Var(R_q) ≤ 1 − C(q) + δ reduces to the simpler 1 − C(q).
-
Sharp versus relaxed form: The sharp certificate replaces the pairwise sum with a weighted multi-firm transport dispersion, C^♯(q) = [L⁻¹ √D_q(C₁,…,C_n) − τ_q]₊², with aggregate slack radius τ_q = (Σ_i q_i τ_i²)^{1/2}. The sharp form improves on the pairwise relaxation in two independent ways: the multi-firm dispersion is at least the weighted pairwise sum of squared Wasserstein distances, and slack is applied once at the aggregate level rather than as τ_i + τ_j per pair. Under a common radius τ, the second effect alone replaces a per-pair deduction of 2τ with a single τ. The pairwise credit remains the decision form because the sharp object requires solving an inner free-centre barycentre problem at each candidate q.
-
Two-asset nesting and envelope exactness: For n = 2, the construction recovers the pairwise covariance envelope of Gawronsky and Huang (2026b). If a coherent joint law attains the dispersion minimum, Σ_i q_i v_i − D_q(P₁,…,P_n) is the greatest systematic portfolio variance attainable across coherent joint laws.
-
Existence and convexity: Any nonempty compact feasible subset of the long-only simplex admits a minimizer of the certified-variance objective. The standardized objective 1 − C(q) is convex along mixtures of normalized risk weights whenever the pairwise floor matrix ℓ is conditionally negative definite — a condition checkable directly from the reported W₂ matrix via Schoenberg's criterion, without returns, and the paper states it holds for the frozen matrix used.
-
Empirical ranking: In the 52-firm panel, the news-only allocation (the zero-slack implementation, formed from observed W₂ geometry without using cross-asset return covariance) lies between the 0.690th and 1.330rd in-sample variance percentiles; equal risk weighting — the inverse-volatility capital benchmark in normalized risk-weight coordinates — lies between the 21.060th and 28.630th percentiles. Across four prespecified capped portfolio populations, between 0.690% and 1.330% of feasible portfolios have variance no greater than the news-only allocation, versus between 21.060% and 28.630% for equal risk weights.
-
Evaluation scope: The authors state that these rankings are descriptive and in-sample, illustrating the allocation implied by the maintained model rather than forecast out-of-sample performance. The lower in-sample variance ranking relative to equal risk also appears across the reported frozen language-model representations.
Methodology in Plain English
-
Represent each firm as a distribution, not a vector. Instead of pooling a firm's news articles into one embedding, the authors keep the whole cloud of article embeddings as an empirical probability distribution. This preserves within-firm heterogeneity that a single pooled vector would suppress.
-
Measure how different two firms' information is. They use the quadratic Wasserstein distance W₂, the minimum root-mean-square displacement needed to rearrange one firm's article cloud to match the other's.
-
Bridge from information to risk via three maintained assumptions. A common carrier map that is L-antilipschitz stops economically distinct information states from collapsing into identical systematic exposures; firm-specific slack radii τ_i allow bounded departures from that common map; and one coherent joint law ensures all pairwise exposure relations can coexist in a single portfolio. These links are maintained restrictions, not facts identified by the text.
-
Convert separation into a variance credit. Observed pairwise distances become lower bounds on exposure separation. A weighted Hilbert-space polarization identity, taken under the coherent joint law, converts those separations into a deduction from the perfect-positive-dependence variance benchmark.
-
Choose the portfolio by minimizing the certified bound. The investor minimizes CV(x) = A(x)² {1 − C(q(x))} over a compact long-only feasible set. No expected-return input enters the objective, constraints, or identifying argument.
-
Keep construction separate from evaluation. The allocation is built from W₂ geometry alone and evaluated only afterward on the full-sample standardized covariance matrix, across four prespecified capped portfolio populations.
Why This Matters
Impact on research. The paper offers a way to obtain a one-sided, certified statement about portfolio risk without estimating every covariance entry, and it adds a portfolio-level aggregation step to the pairwise and cross-sectional distributional-geometry results it builds on. It is explicit that pairwise optimal couplings need not be the pairwise marginals of any single joint law, so separately attainable covariance envelopes cannot simply be stacked — a compatibility requirement that earlier decomposition-based approaches do not enforce. It also connects the textual-finance literature, whose targets are typically expected returns, factors, or pricing, to a coherent bound on cross-asset portfolio risk.
Real-world applications:
- Portfolio construction in markets with short return histories or many assets, where the sample covariance is unstable or singular.
- Diversification monitoring: firms whose information distributions are far apart cannot have perfectly aligned latent risks, which can be used as a screening signal.
- Risk budgeting without a full covariance estimate, using only marginal volatility scales and observed pairwise information distances.
- Strategies built from news or disclosure text where embeddings are already available, allowing the risk inputs to be reused from an existing NLP pipeline.
Industry relevance. The decision rule is a convex program under a condition checkable directly from the observed distance matrix, and the objective needs only marginal volatility scales — inputs a portfolio manager already has. The framework does not replace a covariance estimate with a point estimate; it supplies a bound that can be evaluated at feasible weights and minimized, which fits risk-cap and constraint-driven mandates. The authors caution that the reported rankings are in-sample and descriptive.
Future Directions
- Out-of-sample validation. The reported rankings are in-sample and descriptive; testing the allocation on held-out data is the natural next step and is explicitly not claimed by the paper.
- Calibrating the maintained restrictions. The carrier constant L, the slack radii τ_i, and the residual budget δ are declared inputs rather than estimated parameters; how to set or discipline them from data is left open.
- Numerical evaluation of the sharp certificate. The empirical exercise reports the pairwise relaxation; comparing it against the sharp multi-firm dispersion C^♯(q) on the same data would quantify the tightening in practice.
- Scaling beyond the 52-firm panel and the frozen representations. Whether the certificate and its convexity condition hold, and how tightly they bind, in larger panels and with other language models is not reported.
Target Audience
Researchers in quantitative finance and financial econometrics working on robust or covariance-free portfolio choice; specialists in optimal transport and distributional methods applied to economics and finance; and NLP-for-finance practitioners interested in turning text representations into risk statements. It also suits graduate students and advanced practitioners with a background in portfolio theory, measure-theoretic probability, and optimal transport, since the derivations rely on Hilbert-space identities and Wasserstein stability arguments. Readers seeking empirical out-of-sample performance evidence will not find it here.
Authors’ abstract
Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-level distribution-valued characteristics can instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is convex under a checkable condition and requires marginal volatility scales but no cross-asset return covariances. With zero firm-specific slack, the common-map scale changes the certified variance reduction but not the normalized allocation, which depends only on observed information geometry. In a 52-firm panel from 2018-2022, an allocation constructed from Qwen3-Embedding-8B news representations lies between the 0.69th and 1.33rd in-sample variance percentiles across four prespecified capped portfolio populations; equal risk weighting lies between the 21.1st and 28.6th percentiles. The lower in-sample variance ranking relative to equal risk also appears across the reported frozen language-model representations. The framework therefore distribution-valued firm information into a coherent risk bound and an implementable allocation rule constructed without cross-asset return covariances.