Skip to content
AI.info

Research

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations

Overview Research area: Spatial asset pricing and financial econometrics, combined with text-based finance and optimal transport theory. The paper sits at the intersection of spatial autoregressive mo

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations
arXiv
2608.29669
Published
2026-08-30
Authors
Marcus Gawronsky, Chun-Sung Huang

AI summary

Overview

Research area: Spatial asset pricing and financial econometrics, combined with text-based finance and optimal transport theory. The paper sits at the intersection of spatial autoregressive models, Wasserstein barycentric geometry, and language-model representations of firm news.

Technical level: Advanced. The paper develops formal definitions, lemmas, and theorems (simplex-constrained optimal transport, quadratic adjustment equilibria, spatial closure, stability conditions, Neumann-series reduced forms) and estimates a quasi-maximum-likelihood model on a frozen interaction field.

One-sentence scope: The paper turns the inter-firm interaction matrix of a spatial factor model from a researcher-supplied input into an estimated object, built by reconstructing each firm's news-embedding distribution from those of other firms via target-anchored Wasserstein barycentric weights, then tests whether returns exhibit dependence along that structure out of sample.

What This Paper Is About

Spatial factor models explain cross-sectional return dependence through an interaction matrix W that records which firms are peers and how much each peer matters, but researchers normally supply that W from geography, industry, supply chains, or news co-mentions, leaving the economically central object — the structure of relations among firms — as an assumption rather than a measurement. This paper infers that structure from firms' information environments: each firm is represented by the full distribution (not the average) of its news-article embeddings, and for each target firm a target-anchored Wasserstein barycentric reconstruction selects the nonnegative, unit-sum combination of other firms whose aligned information footprints jointly reconstruct the target's own. The resulting directed "barycentric interaction field" is then frozen before the return window and inserted into a quadratic exposure-adjustment model, whose solved closure produces a spatial lag with coefficient ρ that is estimated on held-out returns.

Key Contributions

  1. A measurement construction for the peer matrix. The paper defines a row-stochastic, zero-diagonal interaction matrix (Definition 1) and constructs its weights (Definition 2) through two stages: optimal transport aligns each target firm's article cloud with every candidate peer's cloud, then a leave-one-out simplex problem chooses the weights that jointly reconstruct the target. The field requires no kernel bandwidth and is directed even though pairwise Wasserstein distance is symmetric, so firm j can be essential for reconstructing firm i while i contributes little to reconstructing j.

  2. An economic interpretation of the field. A quadratic adjustment problem (Definition 3) balances departure from a firm's stand-alone exposure ξ_i against misalignment with peer exposures. Lemma 1 and Theorem 1 show that equilibrium yields the vector-valued spatial autoregression B = ρWB + (1 − ρ)ξ with ρ = λ/(1 + λ), and Equation (8) inverts the mapping via λ = ρ/(1 − ρ). Theorem 2 gives the unique reduced form B = (1 − ρ)(I − ρW)⁻¹ξ under the stability condition ‖ρW‖ < 1.

  3. A joint two-field closure. Definition 4 and Theorem 3 allow the barycentric field and a persistent news co-mention field to enter the same adjustment problem as separate channels, with ρ_B + ρ_N = λ/(1 + λ) as the total peer-adjustment weight and λ_k = ρ_k/(1 − ρ_B − ρ_N) recovering each channel's adjustment index (Proposition 1, Equation 12).

  4. Out-of-sample evidence from language-model representations. Using fields built from 2018–2022 news, embedded with Qwen3-Embedding-8B, for 100 large-cap U.S. firms from a Nasdaq-100-based universe, and frozen before 2023–2026 returns, the paper reports that the field organizes return dependence beyond standard factors and beyond explicit news links.

Main Findings

  • The field organizes return dependence beyond standard factors. Conditional on firm-specific loadings on the Fama–French five factors and momentum, the spatial coefficient is ρ̂ = 0.686596, and the field raises the held-out mean Gaussian quasi-log score relative to an otherwise identical factor-only model, with paired intervals above zero. The spatial coefficient is stable across factor sets and the gain survives firm-specific innovation variances.

  • The gain is in residual covariance, not factor loadings. Because the spatial model leaves reduced-form factor betas exactly equal to firm-by-firm least squares, the paper states the improvement is a statement about the joint distribution of residual returns rather than about factor loadings.

  • Joint reconstruction is what carries the content. Under the same factor-conditioned specification, a field converting the same pairwise Wasserstein distances into proximity weights (the RBF comparator) recovers less than a third of the barycentric field's held-out gain, and weighting the barycentric peers equally recovers less than half. A centroid version of the same reconstruction, which collapses each firm to one average position, recovers most of the gain. Most of the return content therefore comes from changing the economic question from pairwise similarity to target-specific joint representability, with the full within-firm distribution adding a smaller, separately measurable increment.

  • Language-model geometry and explicit news links are distinct channels. Under the primary factor-conditioned specification, the barycentric field and the persistent news co-mention field select partly overlapping but different peers, and each adds conditional fit once the other is present. A raw-return joint model gives the same channel comparison as an appendix corroboration.

  • Robustness to the measurement design. Linear and quadratic transport produce peer-return signals correlating at 0.999775 and predictively equivalent held-out scores, and a dimension ladder shows target reconstruction maps different article-level alignments into nearly identical operators. The barycentric field fits better in sample than pairwise proximity under every encoder examined, across model widths, sizes, families, and training vintages, and a pre-period EttaX V0 encoder preserves the positive W2 gain and its advantage over RBF under the primary held-out design (Appendix D.5).

  • Nesting and dispersion results. At ρ = 0 the model exactly recovers B = ξ, the exposure object of the pairwise covariance restriction in Gawronsky and Huang (2024). Under a nonnegative row-stochastic field, adjustment cannot increase stationary-weighted exposure dispersion and retains at least {(1 − ρ)/(1 + ρ)}² of it (Theorem 5 in Appendix A).

Methodology in Plain English

The approach has four linked steps.

First, each news article about a firm is turned into a position in a high-dimensional embedding space using a language model, so every firm ends up with a cloud of points — an information footprint — rather than a single average location. The paper emphasizes that a diversified firm spreads across many positions, and two firms with the same average embedding can occupy very different footprints.

Second, for each target firm the method aligns its articles with those of every other firm through optimal transport (minimizing average squared displacement), holding those alignments fixed. It then asks: what nonnegative weights that sum to one, over the aligned peer clouds, reconstruct the target's own cloud as closely as possible? Those weights become the target's row of the interaction matrix. Repeating this for every firm produces the barycentric field W^flat, whose rows automatically satisfy the nonnegativity, unit-sum, and zero-diagonal requirements without a separate normalization step or a bandwidth choice.

Third, the field is given economic content. Each firm has a stand-alone exposure from its own information; peer adjustment trades off departing from that exposure against disagreeing with peer exposures, with intensity λ. Solving that quadratic problem produces a spatial autoregression in exposures, and the coefficient ρ = λ/(1 + λ) measures the relative weight of peer alignment. Projecting onto a scalar risk direction carries the same field and coefficient to returns, giving the estimating equation r_t = ρWr_t + x_t.

Fourth, the field is built only from 2018–2022 news, frozen, and then evaluated on 815 trading days of 2023–2026 returns, so the interaction structure is predetermined relative to the outcomes it is used to explain — the conditioning structure that classical spatial-autoregressive inference requires. Comparisons isolate one design margin at a time: equal weights on the selected peers (does the weighting matter?), a centroid representation (does within-firm dispersion matter?), an RBF proximity rule on the same Wasserstein distances (does joint reconstruction matter?), persistent news co-mentions (does explicit shared coverage explain it?), and an industry/sector benchmark.

Why This Matters

Impact on research. For spatial finance, the paper moves the analysis one level upstream: the interaction field becomes an estimated object with an explicit construction and an economic interpretation rather than a researcher's choice. For text-based finance, it repositions language-model representations as a measurement instrument for inter-firm structure rather than as a predictor of returns or a source of factors. For the news-network literature, it shows distributional similarity in news content captures peer relations that explicit co-mentions only partly reflect. Because the construction needs only a firm-level text corpus, an embedding model, and a return panel, the paper argues it applies wherever such data exist.

Real-world applications (as identified by the authors):

  • Covariance estimation — supplying a predetermined interaction structure for residual covariance modeling.
  • Peer benchmarking — building firm-specific peer baskets from information footprints rather than industry codes.
  • Spillover designs — providing a peer structure fixed in advance of the outcome window, avoiding mechanical reflection from estimating W on the outcomes that enter the spatial lag.
  • Portfolio and index construction that requires a defensible, reproducible peer weighting scheme.

The authors also note the method can serve allocation and covariance-bound applications studied separately in Gawronsky and Huang (2024, 2026).

Industry relevance. The method produces a directed, row-stochastic, kernel-free peer matrix of the type factor-model and risk-model providers already consume, with the practical advantage that it is generated from news text plus an embedding model plus a return panel rather than from proprietary supply-chain or analyst data. The finding that the gain appears in residual covariance rather than factor betas is directly relevant to risk-model vendors, whose factor exposures and residual correlation matrices are estimated separately.

Future Directions

  • Separating the pseudo-true coefficient from the structural one. The paper notes that ρ̂ is the pseudo-true coefficient of the concentrated working quasi-likelihood and coincides with the structural coefficient only when the population quasi-score has its unique zero at the structural value despite the richer innovation structure; the return-bridge restriction is an additional maintained condition, so identifying when that holds is left open.
  • Extending the measurement beyond the current universe and window. The evidence covers 100 large-cap U.S. firms from a Nasdaq-100-based universe, 2018–2022 construction news, and 815 trading days of 2023–2026 returns; whether the barycentric–proximity ordering holds for smaller firms, other markets, other asset classes, or longer horizons is not established here.
  • Widening the field set. The two-field adjustment problem admits several admissible fields at once, and the paper demonstrates only the barycentric and persistent co-mention channels; other information objects (filings, transcripts, supply-chain disclosures) could enter as additional channels with their own ρ_k.
  • Tuning and design choices. Squared reconstruction loss, simplex normalization, equal cloud size, and the deterministic rule for selecting among tied optimal assignments are described as maintained design choices, which invites analysis of how sensitive the field is to each.

Target Audience

This paper is aimed at financial econometrics and asset-pricing researchers working on spatial or network models of return dependence, and at researchers using text and language-model representations for measurement rather than prediction. It will also interest quantitative risk-model and factor-model practitioners who construct peer or covariance matrices, and readers with graduate-level preparation in optimal transport, spatial autoregression, and quasi-maximum-likelihood estimation, since the core contributions are formal definitions and theorems alongside the empirical evidence. Readers seeking a light introduction to spatial asset pricing or to embedding-based finance will find the mathematical development demanding; the empirical design, by contrast, is legible to anyone familiar with out-of-sample factor-model comparison.

Authors’ abstract

Spatial asset-pricing models take the structure of inter-firm interaction as given. We infer that structure from firms' information environments using language-model representations. Each firm is represented as a distribution of news-article embeddings, and a target-anchored Wasserstein barycentric reconstruction selects, for every firm, the weighted combination of other firms whose information footprints jointly reconstruct its own. The resulting directed peer field enters a quadratic exposure-adjustment model in which the spatial coefficient indexes alignment with information peers relative to stand-alone exposure. Using fields built from 2018-2022 news and frozen before 2023-2026 returns, we find that the constructed field organizes cross-sectional return dependence beyond the Fama-French five factors and momentum and raises the held-out mean Gaussian quasi-log score relative to a matched factor-only model. Because factor betas are unchanged, the gain lies in residual covariance. The field outperforms pairwise distance weighting and equal weighting of the same peers, and remains incrementally informative beside persistent news co-mentions under the primary factor-conditioned specification. Linear and quadratic transport generate nearly identical peer-return signals and equivalent held-out predictive performance. The barycentric-proximity ordering persists across alternative embedding models, and a pre-period encoder preserves the held-out advantage under the primary specification. Language-model representations thus serve as a measurement instrument for latent inter-firm information structure in capital markets.

Read the original paper