Research
Diving into Kronecker Adapters: Component Design Matters
Overview Research area: Parameter-Efficient Fine-Tuning (PEFT) of large pretrained models, specifically adapter methods that use the Kronecker product (KronA, LoKr, MoKA, MoKA-MoE) to construct weight

- arXiv
- 2602.01267
- Published
- 2026-02-01
- Authors
- Jiayu Bai, Danchen Yu, Zhenyu Liao, TianQi Hou, Feng Zhou, Robert C. Qiu, Zenan Ling
AI summary
Overview
- Research area: Parameter-Efficient Fine-Tuning (PEFT) of large pretrained models, specifically adapter methods that use the Kronecker product (KronA, LoKr, MoKA, MoKA-MoE) to construct weight updates.
- Technical level: Advanced. The paper combines a Kronecker-product singular value decomposition analysis, a subspace-alignment theorem, gradient-norm stability analysis, and multi-model empirical benchmarking.
- Scope: A single paper proposing Component Designed Kronecker Adapters (CDKA), a principled way to choose the three component-design hyperparameters (r₁, r₂, r) of Kronecker adapters, plus guidelines for fixed-parameter budgets and a training stabilization scheme.
What This Paper Is About
Kronecker adapters update a frozen pretrained weight W₀ by adding a sum of r Kronecker products, ΔW = Σᵢ B⁽ⁱ⁾ ⊗ A⁽ⁱ⁾, where A⁽ⁱ⁾ has shape r₁ × (d_in/r₂) and B⁽ⁱ⁾ has shape (d_out/r₁) × r₂. Prior work treated the component design (r₁, r₂, r) as a fixed or hand-tuned heuristic, even though these choices control both the parameter count and the attainable rank. This paper shows that component design, not attainable rank alone, governs how well Kronecker adapters align with full fine-tuning and how well they perform, and it turns that insight into a concrete method (CDKA) with practical configuration and stabilization guidelines.
Key Contributions
- Identifies component design as the key capacity factor. The paper shows that the maximum attainable rank of ΔW is r·r₁·r₂ while the parameter count scales as r(r₁/r₂ + r₂/r₁), so attainable rank alone does not predict performance, and different configurations with identical rank behave very differently.
- Provides a theoretical alignment analysis via Kronecker SVD. Using the Kronecker product singular value decomposition, the paper derives how the subspace alignment between Kronecker adapters and the first-step gradient of full fine-tuning depends on r₁, r₂, and r, yielding three design principles.
- Proposes CDKA with parameter-budget-aware guidelines. The paper gives rules for choosing r₁, r₂ and r under a fixed parameter budget, and tunes the boundary condition r* ∈ [2, 8].
- Introduces a training stabilization strategy. A scaling factor λ_{r₁,r₂,r} = α/√(r·r₂) is shown to make the gradient norm independent of component design, preventing the gradient collapse observed when λ ≡ 1.
Main Findings
- Rank is necessary but insufficient. In Table 2, all configurations with attainable rank 8 on GSM8k using LLaMA-2-7B achieve near-identical or worse results than rank 4: (r₁,r₂,r) = (2,2,1) gives 49.93 ± 1.25, (4,2,1) gives 49.58 ± 0.34, (2,4,1) gives 50.45 ± 0.54, and (2,2,2) gives 51.58 ± 0.18. Reaching attainable rank 4096 with (64,64,1) yields 49.00 ± 0.41, i.e., full rank does not help.
- Increasing r₁ degrades performance. Table 3 shows GSM8k scores of 49.93 ± 1.25 (r₁ = 2), 49.58 ± 0.34 (r₁ = 4), 48.80 ± 1.09 (r₁ = 8), and 49.89 ± 0.27 (r₁ = 16), which matches principle 1.
- Increasing r₂ consistently improves performance. Table 4 shows 49.93 ± 1.25 (r₂ = 2), 50.45 ± 0.54 (r₂ = 4), 53.17 ± 0.43 (r₂ = 8), and 53.93 ± 0.46 (r₂ = 16).
- Increasing r helps only up to r*, then fluctuates. Table 5 shows 49.93 ± 1.25 (r = 1), 54.56 ± 1.62 (r = 4), 54.18 ± 0.72 (r = 16), and 54.26 ± 0.69 (r = 64): a large initial gain followed by instability, consistent with the theoretical claim that performance saturates once r ≥ r*.
- Under a fixed parameter budget, symmetric growth is safe. Table 6 shows (2,2,8) = 56.71 ± 0.38, (4,4,8) = 56.58 ± 0.42, (8,8,8) = 56.38 ± 0.90, and (16,16,8) = 56.56 ± 0.64, i.e., raising r₁ and r₂ together keeps performance stable because their ratio (and hence the budget) is unchanged.
- First spend budget on r, then on r₂. Table 7 shows (2,2,8) = 56.71 ± 0.38, (2,16,2) = 55.17 ± 1.04, (2,2,32) = 56.15 ± 0.75, and (2,16,8) = 57.95 ± 0.43, so raising r to 8 and then r₂ to 16 is the best configuration in that comparison.
- Initialization choice matters. Table 8 shows A⁽ⁱ⁾ Kaiming-initialized with B⁽ⁱ⁾ = 0 gives 56.71 ± 0.38, versus 55.12 ± 1.12 for the reverse (A = 0, B Kaiming), 52.59 ± 0.85 for A Kaiming-normal with B = 0, and 51.86 ± 0.53 for A = 0 with B Kaiming-normal.
- Stabilization requires λ ∈ Θ(1/√(r·r₂)). Theorem 3.4 states the CDKA gradient norm is independent of r₁, r₂, and r if and only if λ_{r₁,r₂,r} ∈ Θ(1/√(r·r₂)); the paper uses λ = 16/√(r·r₂) (α = 16), which equalizes gradient norms across configurations where λ ≡ 1 caused collapse.
- NLU results on GLUE with T5-Base. CDKA with 0.41M parameters reaches an average of 87.30, compared with Full fine-tuning at 226M parameters (87.91), LoRA at 3.24M (82.08), PiSSA (84.71), rsLoRA (79.64), LoRA+ (84.95), DoRA (82.57), AdaLoRA at 4.86M (81.62), GoRA at 3.05M (87.96), LoRA-GA (87.77), LoRA-One (88.73), and KronA* at 0.41M (84.38). The paper describes this as near-optimal performance using only 12.5% of the trainable parameters.
- Mathematical reasoning and code generation with LLaMA-2-7B. CDKA reaches 56.71 ± 0.38 on GSM8k and 24.59 ± 2.74 on HumanEval, versus Full fine-tuning (59.36 ± 0.85 / 35.31 ± 2.13), KronA* (49.00 ± 0.41 / 17.21 ± 2.01), GoRA (54.04 ± 0.22 / 24.80 ± 1.04), LoRA-One* (55.40 ± 0.37 / 20.73 ± 1.00), and DoRA (53.07 ± 0.75 / 19.75 ± 0.41).
- LLaMA-3.1-8B and Qwen results. On GSM8k with LLaMA-3.1-8B, CDKA scores 73.74 ± 0.42, above Full (73.69 ± 0.28), GoRA (72.91 ± 0.76), KronA* (68.11 ± 0.38), and LoRA (67.78 ± 1.25). On MetaMathQA with Qwen-3-0.6B and Qwen-3-8B, CDKA reaches 65.83 ± 0.37 and 86.56 ± 0.20, versus LoRA-One* at 64.77 ± 0.59 and 85.98 ± 0.32, and KronA* at 60.85 ± 0.13 and 85.06 ± 0.67.
- CV coverage claimed but numbers not in the provided text. The paper states that experiments use CLIP-ViT-B/16 across seven image classification tasks and reports state-of-the-art image classification performance, but the truncated content contains no image classification result table, so no figures are quoted here.
Methodology in Plain English
The authors start from a general Kronecker adapter in which every weight update is a sum of r terms, each a Kronecker product of a small matrix A⁽ⁱ⁾ and another small matrix B⁽ⁱ⁾. Two hyperparameters, r₁ and r₂, set the shapes of those small matrices, and r sets how many such terms are summed. LoRA is the special case r₁ = r₂ = 1, and KronA is the special case r₁ = r₂ with r = 1.
They then ask how these three numbers affect learning. First, they show algebraically that the parameter count grows like r(r₁/r₂ + r₂/r₁) while the maximum achievable rank grows like r·r₁·r₂, which explains why Kronecker adapters can reach high rank cheaply. Second, they set up a simplified linear regression problem, the same style of analysis used for LoRA theory, and track the first-step gradient of full fine-tuning versus the adapter's update using the Kronecker product singular value decomposition. The resulting theorem bounds how much of the adapter's leading singular subspace can lie outside the leading singular subspace of the full fine-tuning gradient, and the bound depends explicitly on r₁, r₂, and r, and on quantities such as the initialization scale α, the condition number of the reshaped gradient, and the number of gradient steps t*. Reading off which parameter choices tighten the bound gives the three design principles: keep r₁ small, make r₂ large, and grow r only until it reaches the effective rank r*.
To make r₁, r₂, and r behave consistently during training, they derive a scaling factor λ that makes the gradient norm independent of component design, then set λ = α/√(r·r₂) with α = 16. They also check four initialization schemes and keep A⁽ⁱ⁾ randomly initialized with B⁽ⁱ⁾ = 0, mirroring standard LoRA practice.
Finally, they validate the principles empirically by fine-tuning LLaMA-2-7B on a 100K subset of MetaMathQA and testing on GSM8k while varying one hyperparameter at a time, then benchmark CDKA against LoRA and its variants on GLUE (T5-Base), GSM8k and HumanEval (LLaMA-2-7B), GSM8k (LLaMA-3.1-8B), MetaMathQA (Qwen-3-0.6B and Qwen-3-8B), and image classification (CLIP-ViT-B/16). Results are reported as mean and standard deviation over three random seeds, under the same number of training epochs.
Why This Matters
Impact on research. The paper reframes adapter design: instead of asking "how much rank can this adapter express," it asks "how is that rank distributed across components." It shows that current Kronecker adapters, despite being able to reach far higher attainable rank than LoRA, were underperforming LoRA precisely because of poor component choices. A single tunable structure with no constraints on r₁, r₂, or r generalizes both LoRA and KronA, so the design principles and the λ stabilization rule carry over to any future Kronecker-style adapter. The GitHub code is at https://github.com/rainstonee/CDKA.
Real-world applications.
- Fine-tuning instruction models on mathematical reasoning and code generation, where CDKA's GSM8k and HumanEval results are close to full fine-tuning at a fraction of the trainable parameters.
- Deploying many task-specific adapters on a single pretrained backbone for NLP understanding tasks, since CDKA matches much larger adapters with 0.41M trainable parameters on GLUE.
- Multimodal or vision adaptation, since the paper applies the same approach to CLIP-ViT-B/16 across seven image classification tasks.
- Serving very large or small models under tight memory budgets (the paper evaluates both Qwen-3-0.6B and Qwen-3-8B).
Industry relevance. Serving is the binding constraint for many teams: cheaper adapters mean more tasks per accelerator and lower fine-tuning bills. The paper directly addresses the practical problem that prior Kronecker methods needed manual tuning of component shapes, by giving explicit guidelines (r₁ ∈ [2, 4], r* ∈ [2, 8] as a boundary for r versus r₂) and a closed-form scaling factor that keeps training stable across configurations. One author is affiliated with Huawei, which suggests industrial interest in deployment.
Future Directions
- Theory beyond the linear setting. The alignment analysis uses a linear loss with isotropic sub-Gaussian inputs; extending it to multi-layer, non-linear networks would test whether the principles hold at scale.
- Automatic component design. MoKA uses heterogeneous (r₁, r₂) across components but with manual adjustment; the paper's principles suggest an automatic allocation rule, analogous to AdaLoRA's adaptive rank allocation, that is still open.
- Systematic treatment of r*. The paper empirically suggests r* ∈ [2, 8] as the boundary, but it does not give a way to estimate r* from data or model properties before training.
- Broader modality coverage. Image classification results are claimed but not quantified in the provided text; a fuller comparison on vision and other modalities, and against newer PEFT baselines, would clarify how general the component-design principles are.
Target Audience
Researchers and engineers working on parameter-efficient fine-tuning of large language models and vision models, especially those already using LoRA variants or Kronecker-based adapters such as KronA, LoKr, MoKA, and MoKA-MoE. It is also relevant to practitioners who must pick adapter hyperparameters under a fixed memory or parameter budget, and to theory-oriented readers interested in subspace alignment between adapters and full fine-tuning. A background in linear algebra (Kronecker products, SVD, matrix norms) and familiarity with LoRA is assumed.
Correspondence: lingzenan@hust.edu.cn. Affiliations: Huazhong University of Science and Technology; Huawei; Renmin University of China. arXiv:2602.01267v3 [cs.LG] 08 Aug 2026.
Authors’ abstract
Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component structures. However, existing work largely treats the component structure as a fixed or heuristic design choice, leaving the dimensions and number of Kronecker components underexplored. In this paper, we identify component structure as a key factor governing the capacity of Kronecker adapters. We perform a fine-grained analysis of both the dimensions and number of Kronecker components. In particular, we show that the alignment between Kronecker adapters and full fine-tuning depends on component configurations. Guided by these insights, we propose Component Designed Kronecker Adapters (CDKA). We further provide parameter-budget-aware configuration guidelines and a tailored training stabilization strategy for practical deployment. Experiments across various architectures and modalities demonstrate the effectiveness of CDKA. Code is available at https://github.com/rainstonee/CDKA.