Skip to content
AI.info

Research

On the Expressive Power of Permutation-Equivariant Weight-Space Networks

On the Expressive Power of Permutation-Equivariant Weight-Space Networks Overview Research area: Machine learning theory — specifically the expressive power (approximation capability) of weight-space

On the Expressive Power of Permutation-Equivariant Weight-Space Networks
arXiv
2602.01083
Published
2026-02-01
Authors
Adir Dayan, Yam Eitan, Haggai Maron

AI summary

On the Expressive Power of Permutation-Equivariant Weight-Space Networks

Overview

Research area: Machine learning theory — specifically the expressive power (approximation capability) of weight-space neural networks, which take the parameters of other neural networks as input. The paper sits at the intersection of Geometric Deep Learning, symmetry/equivariance theory, and weight-space learning.

Technical level: Advanced. The paper is a theoretical treatment built on group representations, permutation equivariance, and universal approximation arguments, with one empirical validation on a model-editing benchmark.

Scope: A unified expressivity theory for permutation-equivariant weight-space networks operating on MLP weights, covering four approximation settings, plus a theory-driven architectural modification that improves prior state of the art by up to 34% on INR editing.

Authors and affiliations: Adir Dayan and Yam Eitan (Technion – Israel Institute of Technology, equal contribution), Haggai Maron (Technion and NVIDIA Research). Posted as arXiv:2602.01083v2 [cs.LG], licensed CC BY 4.0. Keywords listed: Machine Learning, ICML.

What This Paper Is About

Weight-space networks operate directly on the parameters of trained neural networks, and the leading ones are designed to be permutation-equivariant: permuting neurons in a hidden layer changes the weights but not the function the network computes, so these architectures build that symmetry in. Symmetry-preserving designs restrict the hypothesis space, which raises the question of whether they lose approximation power. The paper asks what these networks can and cannot approximate, and it notes that the question is subtler than in other symmetry domains because weight-space learning targets maps on both weight space and function space.

Key Contributions

  1. Expressive equivalence of architectures. The authors prove that all prominent permutation-equivariant weight-space networks for MLPs are equally expressive: Deep Weight Space (DWS) networks, the neuron-permutation and hidden-neuron-permutation variants of Neural Functional Networks (NP-NFN and HNP-NFN), Graph Meta-Networks (GMNs), and Neural Graph GNNs (NG-GNNs). Neural Functional Transformers (NFTs) are the sole exception in full generality, but match the rest under a general-position assumption on the input weights.

  2. An approximation framework with four settings. The paper identifies four natural approximation settings — function-space functionals, permutation-invariant functionals, function-space operators, and permutation-equivariant operators — and defines a formal notion of approximation for each.

  3. A universality characterization. For each setting the authors establish when universality holds, prove universality under natural general-position assumptions, and identify the regimes where universality fails.

  4. Empirical gains from theory. Guided by the analysis, the authors propose a simple modification to existing weight-space models and report up to a 34% improvement over prior state of the art on a standard model-editing task.

Main Findings

  • All major equivariant weight-space networks are equally expressive. For any compact set K of weights, the classes of invariant maps and equivariant operators approximated by DWS, NP-NFN, HNP-NFN, GMN, and NG-GNN are identical (Theorem 5.2). The proof works by explicitly approximating the base layers of one network with those of another.

  • NFTs are the exception, but only off the general-position set. There exists a compact weight set on which the maps NFT can approximate differ from those of the other architectures (Proposition 5.3). On any compact set of weights outside the bias exclusion set — that is, where all neuron biases within each hidden layer are pairwise distinct — NFT becomes equivalent to the others.

  • The exclusion set is measure zero and natural. The bias exclusion set contains only degenerate configurations and has Lebesgue measure zero, so it is unlikely to arise under random initialization or stochastic training. The authors validate the general-position assumption empirically on standard pretrained networks (Appendix G).

  • Function-space functionals are universally approximable. Every continuous map from functions to output vectors can be approximated by invariant weight-space networks on any compact weight set (Theorem 6.1). The proof uses DWS's ability to simulate a forward pass, a separation property, and a separation-to-approximation result.

  • Permutation-invariant functionals are not universal in full generality. There exists a compact weight set and a permutation-invariant map that invariant weight-space networks cannot approximate (Proposition 6.2). The counterexample uses two binary weight configurations whose second-layer weight matrices have different ranks and whose induced neural graphs are indistinguishable by the Weisfeiler–Leman test.

  • Universality for permutation-invariant functionals returns under general position. On any compact set of weights outside the exclusion set, any permutation-invariant functional can be approximated by permutation-invariant weight-space networks (Theorem 6.3). The construction sorts neurons by bias values to build a continuous canonization map, approximates the ranking with DWS layers (which subsume DeepSets primitives), and composes with an MLP head.

  • Function-space operators are not universally approximable for a fixed input architecture. For any fixed ReLU architecture, there exists a family of natural continuous function-space operators that cannot be approximated by permutation-equivariant weight-space networks defined over that architecture (Proposition 7.1, informal). The intuition is that networks constrained to output weights of the same architecture cannot produce outputs of greater geometric complexity, and ReLU MLPs of fixed size have a bounded number of linear regions.

  • Larger input architectures restore universality for function-space operators. Given a continuous function-space operator and a compact function set, a sufficiently large architecture exists such that the operator can be approximated on compact weight sets in general position whose realized functions approximate those in the function set (Theorem 7.2, informal).

  • Universality may need fewer layer types than expected. The construction for the permutation-invariant functional setting uses only a restricted subset of the DWS operations, suggesting the network can be simplified without loss of expressive power.

  • Low-precision settings are a practical risk zone. When weights or biases are heavily quantized (low-bit or binary networks), inputs are more likely to fall inside the exclusion set, where full universality for invariant weight functionals may fail.

Methodology in Plain English

The authors take a theoretical route. First, they fix notation for MLP weight spaces, define the neuron-permutation group and its representation on weights, and define what it means for a map to be invariant or equivariant. They then define a notion of approximation for each of four target types: vector outputs that depend only on the realized function, vector outputs that are permutation-invariant but may depend on parameterization, function-to-function maps, and weight-to-weight maps.

To compare architectures, they define, for each network family, the set of maps it can approximate on a compact domain, and prove these sets coincide by showing that each network's building blocks can be approximated by another's. For universality results they use a general-position assumption — distinct biases per hidden layer — which lets them build a continuous canonization map that sorts neurons by bias and sends every permutation of a weight configuration to a single canonical representative. Since permutation invariance implies the target map factors through that canonical form, universal approximation reduces to approximating a continuous function on a compact Euclidean set. To show limits, they construct adversarial weight configurations that message-passing-based networks cannot distinguish, while a simple invariant quantity (matrix rank) can. The empirical part applies the theory: instead of predicting one output network, the model predicts several and ensembles them.

Why This Matters

Impact on research. The paper converts a scattered set of partial capability results into a single expressivity landscape. It tells theorists and practitioners which of the popular equivariant weight-space architectures are interchangeable, which one (NFT) is different and under what condition it is not, and precisely which target types are and are not universally approximable. It also reframes "more architecture" as a lever: for function-space operators, enlarging the input architecture is what restores universality.

Real-world applications (settings listed in the paper's Figure 2, with cited prior work):

  • Model accuracy prediction and INR classification — function-space functionals over trained networks.
  • Image and 3D scene model editing — function-space operators, including zoom-out style transformations of implicit neural representations and NeRFs.
  • Domain adaptation at the function level — mapping a function at a global minimum of one loss to a global minimum of a loss incorporating additional data.
  • Pruning mask prediction and gradient prediction for meta-optimization — permutation-equivariant weight-to-weight operators.

Industry relevance. Weight-space methods matter for anyone using pretrained models as reusable assets. The paper's theory-driven modification delivers up to a 34% improvement over prior state of the art on INR editing without a new architecture, and its warning about quantization is directly relevant to low-bit and binary network pipelines, where degeneracies in the exclusion set become likely.

Future Directions

  • Extending the theory beyond MLPs. The paper notes that universality results analogous to Theorem 6.3 can be obtained for weight-space architectures operating on transformer parameters via minor modifications of the proof, with details in Appendix H — suggesting a broader scope for the framework.

  • Alternative exclusion sets. The bias-based exclusion set is described as one convenient choice, and the authors state the theory extends naturally to several alternatives, discussed in Appendix G.

  • More expressive variants for quantized regimes. Because low-bit and binary networks make degeneracies more likely, the paper suggests that applications involving extreme quantization may benefit from architectural variants with higher expressive power.

  • Larger input architectures for model editing. Most prior model-editing studies use relatively small MLPs, often with only two hidden layers and modest hidden dimensions. The theory indicates that increasing input architectural capacity substantially enhances expressivity, which the authors validate empirically in Section 8.

  • Simplification of existing designs. Since universality in the invariant functional setting was achieved with only a restricted subset of DWS operations, an open question is how far the standard architectures can be pruned without losing expressive power.

  • The paper's closing sentence on future work is cut off in the provided text, so the full list of proposed directions is not reported.

Target Audience

This paper benefits machine learning theorists working on expressivity, universal approximation, and symmetry in neural networks, as well as Geometric Deep Learning researchers who study permutation equivariance in structured domains. It is also relevant to practitioners building weight-space models for model editing, meta-learning, pruning, or accuracy prediction who need to know whether a given equivariant architecture is expressive enough for their task — and to engineers working with quantized or low-bit networks, for whom the exclusion-set caveat is a concrete design consideration. A reader needs comfort with group representations, compactness arguments, and standard approximation-theoretic reasoning; the empirical section alone is accessible to a broader audience.

Authors’ abstract

Weight-space learning studies neural architectures that operate directly on the parameters of other neural networks. Motivated by the growing availability of pretrained models, recent work has demonstrated the effectiveness of weight-space networks across a wide range of tasks. SOTA weight-space networks rely on permutation-equivariant designs to improve generalization. However, this may negatively affect expressive power, warranting theoretical investigation. Importantly, unlike other structured domains, weight-space learning targets maps operating on both weight and function spaces, making expressivity analysis particularly subtle. While a few prior works provide partial expressivity results, a comprehensive characterization is still missing. In this work, we address this gap by developing a systematic theory for expressivity of weight-space networks. We first prove that all prominent permutation-equivariant networks are equivalent in expressive power. We then establish universality in both weight- and function-space settings under mild, natural assumptions on the input weights, and characterize the edge-case regimes where universality no longer holds. Guided by our theoretical results, we show that slight modifications to existing weight-space models yield a 34% improvement over prior SOTA, demonstrating the practical relevance of our framework.

Read the original paper