Research
Weight Space Representation Learning via Neural Field Adaptation
Overview Research area: Machine learning — weight space learning, implicit neural representations (neural fields), low-rank adaptation (LoRA), and diffusion-based generative modeling. Technical level:

- arXiv
- 2512.01759
- Published
- 2025-12-01
- Authors
- Zhuoqian Yang, Mathieu Salzmann, Sabine Süsstrunk
AI summary
Overview
- Research area: Machine learning — weight space learning, implicit neural representations (neural fields), low-rank adaptation (LoRA), and diffusion-based generative modeling.
- Technical level: Intermediate to Advanced. Readers should be comfortable with neural fields/INRs, LoRA, permutation symmetry in networks, and diffusion models.
- Scope: The paper proposes multiplicative LoRA (mLoRA) inside a frozen pre-trained neural field as a way to turn per-instance network weights into structured, semantic data representations, and evaluates those representations on reconstruction, generation, and discriminative tasks over 2D (FFHQ) and 3D (ShapeNet) data.
What This Paper Is About
Neural network weights are normally treated as opaque byproducts of training, but this paper asks whether the weights themselves can act as useful representations of data. The problem is that functionally identical networks can sit arbitrarily far apart in weight space because of symmetries such as neuron permutation and scaling, making the weight distribution multi-modal and hard to model. The authors' goal is to impose enough structure — by adapting a shared pre-trained base neural field with a multiplicative low-rank adaptation — that each instance's weights become a well-behaved, semantically organized representation.
Key Contributions
- Demonstrating that independently optimized neural network weights, when properly constrained, can serve as effective data representations that capture semantic structure.
- Introducing multiplicative LoRA for neural fields, which the authors show provides superior representation quality compared to standard additive LoRA and standalone MLP weight parameterizations.
- Validating weight space representations across diverse tasks — reconstruction, generation, and classification — establishing their viability as a representation paradigm.
- Showing that combining multiplicative LoRA with asymmetric masking removes internal permutation symmetry, producing weights that converge to a linear mode and that generate higher-quality samples than prior weight-space methods.
Main Findings
- Multiplicative beats additive: Across reconstruction, generation, and discriminative evaluation, the six compared representations (MLP, MLP-Asym, LoRA, LoRA-Asym, mLoRA, mLoRA-Asym) rank in a clear progression — LoRA outperforms standalone MLPs, mLoRA outperforms LoRA, and mLoRA gives the best overall discriminative results.
- Weight space structure: In the ShapeNet airplane stability analysis (averaged over 30 instances), weight similarity for MLP and MLP-Asym decreases approximately linearly with perturbation strength, while LoRA-based representations saturate at large perturbation strengths. Applying LoRA alone does not improve linear mode connectivity.
- Asymmetric masking helps structure: The asymmetric mask improves both weight similarity and linear mode connectivity across all parameterizations. mLoRA-Asym retains very high similarity and a very low barrier even under very different initializations, suggesting its weights converge to a linear mode.
- Generation on 2D FFHQ (Table 1, lower is better): mLoRA-Asym achieves FD 0.073, MMD-G 0.039, MMD-P 0.467, versus HyperDiffusion at 0.241, 0.158, 1.887. mLoRA alone reaches 0.100, 0.056, 0.674. LoRA (0.321, 0.169, 2.018) and LoRA-Asym (0.269, 0.157, 1.877) do not beat HyperDiffusion on FD.
- Generation on 3D ShapeNet (Table 2): On the airplane model, mLoRA-Asym achieves mMD 1.89, COV 43.4%, 1-NNA 71.9%, FD 0.011, MMD-G 0.003, MMD-P 0.041, compared with HyperDiffusion at 2.39, 43.6%, 78.2%, 0.027, 0.009, 0.122. On the 10-category model, mLoRA-Asym reaches mMD 5.52, COV 49.6%, 1-NNA 58.6%, FD 0.026, MMD-G 0.004, MMD-P 0.040, versus HyperDiffusion at 8.64, 41.6%, 78.3%, 0.117, 0.023, 0.219. Additive LoRA and LoRA-Asym fail badly on 3D (for example LoRA-Asym mMD 270.6 on airplane, 210.0 on multi).
- Multi-category degradation: HyperDiffusion performs competently on the single-category airplane model but degrades substantially in the multi-category setting. On FFHQ, HyperDiffusion fails to produce recognizable face images, whereas both mLoRA and mLoRA-Asym manage — described as the first successful weight space generation for high-resolution natural image generation, since previous methods were restricted to simpler datasets such as MNIST and CIFAR.
- Discriminative power (Table 3, ShapeNet ten-category, 10 runs): mLoRA leads on clustering ARI (67.1% ± 3.7%), 1-NN accuracy (85.1% ± 1.8%), and logistic regression accuracy (90.0% ± 0.7%). LoRA-Asym reaches 47.4% ± 4.4% ARI and 59.1% ± 10.4% 1-NN, below plain LoRA (56.3% ± 3.3% and 75.2% ± 1.7%). Because mLoRA-Asym has better weight space structure but lower discriminative accuracy than mLoRA (56.5% ± 2.7% ARI, 80.8% ± 1.6% 1-NN, 84.5% ± 1.4% logistic), the authors conclude discriminative power is not directly linked to linear mode connectivity or permutation symmetry.
- Geometry visualization: In t-SNE plots of the ten-category representation, all weight spaces capture some semantic structure, but only multiplicative LoRA weights show clear class separation. In the multi-category perturbation study (20 instances from each of five ShapeNet categories — airplane, car, chair, sofa, table — at 9 perturbation strengths), MLP and MLP-Asym organize as initialization → category → instance, whereas mLoRA and mLoRA-Asym invert this to category → instance → initialization, with additive LoRA in between.
Methodology in Plain English
The authors fit one small neural field per data instance and then use that network's parameters as the representation of the instance. Rather than fitting each network from scratch, they fine-tune a single shared, pre-trained base neural field with low-rank adaptation, keeping the base weights frozen and storing only the small LoRA matrices per instance.
The central modification is the adaptation rule: instead of the standard additive update W + BA, they use an element-wise multiplicative update W ⊙ BA. The authors argue that neural fields build signals additively and therefore entangle features, and that additive LoRA injects more components into that mixture, whereas multiplication scales existing features and preserves channel structure. To break the symmetry that lets the same function be written with different LoRA factors, they apply asymmetric masking: for each layer they randomly freeze sqrt(d_out) entries in each row of the LoRA A matrix, using the same frozen positions across all instances and training runs. For multiplicative LoRA these frozen entries are simply set to zero, which gates off the corresponding rank components; for standalone MLPs and additive LoRA the frozen entries are instead initialized with higher variance N(0, κI) while other weights use N(0, I).
The base model is trained with a variational autodecoder: network parameters and per-instance latent codes are optimized jointly with a reconstruction loss plus an L2 penalty λ_r on the latent codes. For generation, they train diffusion models on the flattened weight representations, using a DDPM forward process with a linear noise schedule from 10⁻⁴ to 2×10⁻² over T timesteps and a diffusion transformer that predicts the added noise; DDIM sampling is used at inference. For LoRA weights they design a hierarchical LoRA layer encoder that treats each (a, b) vector pair as a token, adds rank-level positional encodings, applies multi-head attention with r heads within a layer, then adds layer-level positional encodings before the main transformer.
Evaluation uses FFHQ at 128×128 for 2D and ShapeNet for 3D (a single-category airplane model and a ten-category model). Generation quality is measured with Fréchet distance, MMD with a Gaussian RBF kernel and with a polynomial kernel, plus mMD, COV, and 1-NNA for shapes, using CLIP features for 2D images and a PointNet++ extractor for 3D shapes.
Why This Matters
This work challenges the assumption that weights are unusable, opaque optimization artifacts and shows that a simple change in the adaptation rule plus symmetry breaking can make weight space organized enough to generate from and to classify with. It also links a geometric property of the weight space — how close independently optimized solutions land, and where the linear interpolation barrier sits — to downstream generative quality, which gives a concrete diagnostic for future weight-space methods.
Real-world applications:
- 3D content creation: Generating novel shapes and shape collections from a diffusion model trained over neural field weights, as demonstrated on ShapeNet airplanes and a ten-category setting (airplanes, chairs, tables, and other common objects).
- Signal and asset compression: Since implicit neural representations parameterize data directly, low-rank weight representations are a natural compact per-instance code (the paper notes INRs have been explored for compression in prior work).
- Semantic organization of model or asset repositories: The clustering and 1-NN results on the ten-category dataset suggest weights can be used to retrieve and group items by category without an explicit encoder.
- High-resolution image generation from compact weight codes: The FFHQ results at 128×128 indicate weight-space generation is no longer limited to very simple datasets.
Industry relevance centers on generative pipelines for 2D images and 3D assets, where per-instance low-rank weight codes could be stored, searched, interpolated, and shared more cheaply than full models, and where weight-space diffusion offers an alternative to hypernetworks and conventional encoders.
Future Directions
- The paper's limitations and future research directions are discussed in Section 14 of the supplementary material, but that content is not included in the material available here, so its specific claims cannot be summarized.
- Extending the approach beyond the evaluated settings: higher image resolutions than FFHQ 128×128, more than ten ShapeNet categories, and data modalities other than 2D images and 3D shapes.
- Closing the gap between structure and semantics: mLoRA-Asym has the best weight space geometry but lower classification accuracy than mLoRA, so understanding why structure and discriminability diverge is an open question.
- Reducing reliance on asymmetric masking — for example, finding adaptation mechanisms or architectures that yield well-behaved weight spaces without hand-designed symmetry-breaking constraints, or applying masking more effectively to standalone MLP and additive LoRA parameterizations without the large-variance initialization problem.
- Applying the structured weight space to downstream uses the paper does not evaluate, such as semantic editing or interpolation in weight space (interpolation is studied in the supplementary material).
Target Audience
Researchers and graduate students working on weight space learning, neural fields/implicit neural representations, parameter-efficient fine-tuning, and diffusion-based generation, as well as practitioners building 3D asset generation or compact representation systems who want to understand when network weights can be treated as first-class data.
Authors’ abstract
We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-trained base model and low-rank adaptation (LoRA) can induce structure in weight space. Across reconstruction, generation, and analysis tasks on 2D and 3D data, we find that multiplicative LoRA weights achieve high representation quality while exhibiting distinctiveness and semantic structure. When used with latent diffusion models, multiplicative LoRA weights enable higher-quality generation than existing weight-space methods.