Research
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections
Overview Research area: Dense retrieval / information retrieval (cs.IR), with a theoretical core in statistical learning theory, low-rank matrix estimation, and Riemannian operator geometry. Technical

- arXiv
- 2609.32488
- Published
- 2026-09-26
- Authors
- Maojun Sun, Yancheng Yuan, Jian Huang, Ruijian Han
AI summary
Overview
- Research area: Dense retrieval / information retrieval (cs.IR), with a theoretical core in statistical learning theory, low-rank matrix estimation, and Riemannian operator geometry.
- Technical level: Advanced. The paper uses fixed-rank manifold tangent spaces, Stein's Unbiased Risk Estimate (SURE), noncentral chi-squared selection power, and Rademacher complexity, alongside a full retrieval experimental suite.
- Scope in one sentence: The paper derives an exact bias–variance boundary that says when separate (dual) query and document projections beat a single shared projection in dense retrieval, and turns that boundary into a practical, training-data-only geometry selector called CARS.
What This Paper Is About
When adapting frozen query and document embeddings, a system can apply the same rank-r projection to both sides ("shared") or fit two separate projections ("dual"). Shared projections can only produce positive-semidefinite scoring operators, while dual projections can produce any rank-at-most-r operator, so dual is more expressive but has more parameters to fit from limited training pairs. The paper's goal is to work out exactly when that extra expressive power pays for itself, and to build a selector that decides between the two geometries from training data alone.
Key Contributions
- Exact approximation characterization. The authors characterize the operator classes induced by shared and dual rank-r projections, show that the shared family is a strict subset of the dual family, and derive the exact Frobenius approximation loss imposed by shared projections (Theorem 1), including the exact Shared-minus-Dual approximation gap.
- A local Gaussian bias–variance boundary. They establish a local Gaussian boundary (Theorem 2 and Corollary 1) stating that dual projections have lower risk precisely when squared directional signal exceeds the estimation cost of their additional degrees of freedom, with the extra dimension count k = r(2p − r − 1)/2.
- A Stein-unbiased selection rule with exact power and regret. They derive a SURE-based rule whose selection probability follows a noncentral chi-squared law and whose model-choice regret has a closed form (Theorem 3), then introduce CARS (Cross-fitted Asymmetry Risk Selector) as a cross-fitted version for real embeddings.
- Empirical validation across simulations, full-corpus retrieval, and held-out operator risk. The predicted signal and sample-size effects are tested on controlled two-view retrieval, five datasets with four encoders, and a held-out operator-loss study that evaluates CARS against two fixed-geometry baselines.
Main Findings
- Exact cost of sharing. The minimum squared Frobenius error of a shared rank-r projection decomposes into three terms: skew energy ‖K‖²_F, squared negative eigenvalues of the symmetric part, and discarded positive eigenvalues beyond the largest at most r. The dual minimum is the singular-value tail Σ_{j>r} σ_j(M⋆)², and the gap is the difference.
- The boundary condition. Dual has lower local asymptotic risk than shared if and only if δ² > σ²k, where δ² is the squared directional signal in the dual-only tangent directions, σ² is the noise scale, and k = r(2p − r − 1)/2. The dimension count decomposes as k = r(r − 1)/2 + r(p − r).
- Tangent-space dimensions. dim T_d = 2pr − r², dim T_s = pr − r(r − 1)/2, with T_s ⊂ T_d.
- A coarse global bound is insufficient. The Rademacher bound of Proposition 1 preserves the ordering ℜ_{n,s} ≤ ℜ_{n,d} = B E‖Q_r‖_F ≤ BR/√n, but bounds both families by the same ceiling, so it cannot quantify dual's extra estimation cost.
- SURE corrects plug-in bias. The naive plug-in statistic ‖Π_A Z‖²_F has expectation δ² + σ²k and is biased upward. The SURE rule selects Dual exactly when ‖Π_A Z‖²_F > 2σ²k, yielding Pr(π̂ = d) = Pr{χ²_k(λ) > 2k} with λ = δ²/σ². When δ = 0, the probability of selecting Dual is at most (2/e)^{k/2}.
- CARS expectation identity. Under the independent-half model, E Γ̂_CARS = ‖Δ‖²_F − tr(Σ)/n; when the residual lies in 𝒜 with Δ = n^{−1/2} Π_A H and Σ = σ²I_k, this equals (δ² − σ²k)/n, the same signal-versus-variance comparison as the theoretical boundary.
- Crossover in controlled retrieval. In the rank-8 Gaussian grid, the first positive mean occurs at n = 256 for a 0.8-radian rotation but at n = 64 for a 1.2-radian rotation. The local-Gaussian calibration matches the Theorem 2 risk limits within 0.25% and the Theorem 3 SURE selection probabilities within 0.001 (Appendix Table 2).
- Raw asymmetry is inflated. On real embeddings, the raw plug-in trend is nearly flat across score deciles, whereas the corrected-score trend rises from Shared-favored to Dual-favored held-out fit across 840 operator fits on ten tasks.
- Rotation effect on retrieval. Averaged across five datasets, the Dual-minus-Shared test NDCG@10 gap rises from about .0032 at 0° to .0146 at 90°, a 4.6-fold increase; the abstract describes the mean advantage as more than doubling over the same 0° to 90° range. All five dataset means increase, and on MS MARCO the gap changes sign from about −.0008 to +.0219.
- Sample-size reversal. Shared wins 13 of the 16 displayed cells at n = 32, whereas Dual wins all 32 cells at n = 1024 and 2048. Rank 32 gives slightly higher absolute NDCG@10 than rank 4 for both Shared (0.7242 to 0.7257) and Dual (0.7282 to 0.7289), without enlarging Dual's relative advantage.
- Operator risk shifts with data. Across FEVER, NQ, ArguAna, and SciFact, all 168 comparable sample-size slopes are positive; on NQ, three of the four encoder-mean curves cross from Shared-favored to Dual-favored held-out fit.
- CARS selects well. CARS wins 85/100 encoder–fold comparisons, reaches 90.1% mean geometry-selection accuracy, and reduces mean regret by 49–96% relative to the better fixed-geometry choice. In Table 1, for example, MS MARCO with BGE-base gives Always-Shared regret 0.0680 at 75.6% accuracy, Always-Dual regret 9.5146 at 24.4% accuracy, and CARS regret 0.0207 at 93.8% accuracy.
Methodology in Plain English
The authors start from the observation that shared and dual projections produce mathematically different families of scoring matrices: shared ones must be symmetric positive semidefinite, dual ones can be any low-rank matrix. They first compute exactly how much error the shared restriction costs when the target operator is known, splitting that error into a part that comes from asymmetric (skew) structure and a part that comes from discarding weak or negative eigenvalues.
To study what happens with finite data, they place the problem in a local Gaussian model: the empirical relevance moment is the true rank-r target plus isotropic noise of scale σ/√n, and the target's departure from the shared family is also of order n^{−1/2}. This common scale lets approximation gain and estimation cost be compared directly. Counting the tangent-space dimensions of the two families gives the number of "dual-only" directions k, and the risk difference comes out to exactly δ² − σ²k.
Because δ² is unknown in practice, they build an unbiased estimate using Stein's Unbiased Risk Estimate, which corrects the upward bias in the naive plug-in statistic (the naive estimate is inflated by σ²k expected noise energy in the dual-only directions). The resulting rule selects Dual when a threshold of 2σ²k is crossed, and its behavior is characterized exactly through a noncentral chi-squared distribution.
For real embeddings, where the noise covariance and geometry are unknown, they replace the theoretical quantities with Cross-fitted Asymmetry Risk Selector (CARS): split the training triples into two halves repeatedly, fit each half's shared and dual operators, and score agreement across halves (an inner product) minus disagreement (a scaled squared difference). This score estimates the same signal-minus-variance quantity from data alone.
The empirical program has four questions: whether the predicted crossover appears in controlled two-view retrieval; whether correcting for estimation noise makes observed asymmetry more informative; how mismatch, rank, and training size change full-corpus retrieval quality; and whether held-out operator risk shifts with more data and whether CARS picks the better geometry. Retrieval experiments use frozen E5-base-v2, BGE-base, GTE-base, and Contriever-MSMARCO embeddings, rank-16 adapters trained with 1,024 queries, query rotation through seven angles from 0° to 90° in eight relevance-informed planes, exact inner-product search over complete corpora, and held-out operator loss L_g = ‖M̂_g − M_test‖²_F / ‖M_test‖²_F.
Why This Matters
Impact on research. The paper gives geometry choice in dense retrieval a precise, testable criterion instead of a heuristic. It connects retrieval head design to low-rank matrix estimation, SURE, and manifold tangent-space counting, and it shows why coarse uniform-convergence bounds are too loose to decide the question — a distinction relevant to anyone comparing constrained and unconstrained parameterizations.
Real-world applications:
- Retrieval-augmented generation pipelines, where retrieval quality directly determines the evidence passed to a generator.
- Semantic search over large corpora, where adapting frozen embeddings with a low-rank head is cheaper than retraining encoders.
- Question answering systems, which the paper lists among the systems dense retrieval supplies evidence to.
- Domain adaptation of frozen embeddings with limited labeled query–document triples, where the paper's sample-size results indicate that low-data regimes favor shared projections and higher-data regimes favor dual.
Industry relevance. Practitioners adapting pretrained embeddings face exactly the decision this paper formalizes: whether to fit one projection or two. The rank–sample-size grids give an actionable signal — Shared wins 13 of 16 cells at n = 32, Dual wins all 32 cells at n = 1024 and 2048 — and CARS offers a training-only procedure that reduced held-out regret by 49–96% relative to the better fixed baseline at 90.1% mean selection accuracy across four encoders and five datasets.
Future Directions
- Anisotropic noise. The exact risk boundary assumes local, isotropic Gaussian operator noise; extending risk estimation to anisotropic noise would make the theory more realistic.
- Ranking-metric-aware selection. The authors propose extending risk estimation directly to ranking metrics so that geometry selection more closely reflects retrieval performance rather than operator loss.
- Partial sharing. Selecting how many directions are separately parameterized alongside rank, rather than choosing between fully shared and fully dual projections.
- Data-efficient retrieval-head design and other two-view problems. The framework is proposed as guidance for retrieval-head design and for other two-view representation problems in which the two inputs play different roles.
Target Audience
This paper is most useful to information retrieval and representation learning researchers working on dense retrieval heads and embedding adaptation, and to theoretically inclined readers interested in bias–variance analysis of low-rank matrix families. It also suits applied engineers who must decide between shared and dual projection heads under limited labeled query–document data, and readers of retrieval-augmented generation and semantic search systems who want a principled rule instead of a default. A background in linear algebra, rank-constrained estimation, and basic statistical risk analysis helps; the full derivations in the appendix require comfort with manifold tangent spaces and Stein's risk estimate.
Authors’ abstract
Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclear. We introduce a bias-variance theory for low-rank bilinear scoring. Shared projections induce positive-semidefinite operators, whereas dual projections realize arbitrary low-rank operators. We derive their exact approximation gap and prove a local Gaussian boundary: dual has lower risk exactly when squared directional signal exceeds the estimation cost of its additional degrees of freedom. This boundary motivates the Cross-fitted Asymmetry Risk Selector (CARS), which estimates reproducible directional signal from training pairs; its Gaussian counterpart admits exact selection-power and regret formulas. Guided by the theory, we run retrieval experiments across multiple datasets and embedding models. The mean Dual-minus-Shared NDCG@10 advantage more than doubles as query rotation increases from 0 degrees to 90 degrees. In the rank-sample-size grids, Shared wins 13 of 16 cells at n=32, whereas Dual wins all 32 cells at n=1024 and n=2048. Consistent with this shift, all 168 comparable operator-risk curves move toward Dual as training data grow. Compared to the two fixed-geometry baselines, CARS reduces held-out regret by 49-96% and achieves 90.1% mean geometry-selection accuracy.