Research
On De-Individuated Neurons: Continuous Symmetries Enable Dynamic Topologies
Overview Research area: Neural network architecture design and continual learning (cs.NE), sitting at the intersection of group theory/symmetry methods and adaptive network topologies. Technical level

- arXiv
- 2602.23405
- Published
- 2026-02-26
- Authors
- George Bird
AI summary
Overview
Research area: Neural network architecture design and continual learning (cs.NE), sitting at the intersection of group theory/symmetry methods and adaptive network topologies. Technical level: Advanced — the paper assumes familiarity with group representations, the orthogonal group O(n), singular value decomposition, and independent component analysis. Scope: A single-author theoretical paper (George Bird, University of Manchester) that derives "isotropic" activation functions prescribed by continuous orthogonal symmetry, uses them to diagonalise dense layers into one-to-one connections, and builds a neurogenesis/neurodegeneration procedure on top of that diagonalisation, with a supporting CIFAR-10 experiment on a multilayer perceptron.
Note: the supplied content ends mid-table in Section 3, so the numeric results of Table 1 are only partially legible.
What This Paper Is About
Most activation functions are "elementwise": they apply the same scalar operation independently to each component of a vector, which fixes an arbitrary coordinate basis and treats each component as an individual neuron. The paper argues this "individuation" is what makes it hard to add or remove whole neurons without damaging the network's function, and that it is why individual-connection pruning is more common than whole-neuron removal. The goal is to redefine activation functions so that no particular basis is distinguished, then exploit that basis freedom to rewrite a dense layer as a one-to-one diagonal mapping whose neurons can be added, buffered or deleted while the network's output is preserved or very closely approximated.
Key Contributions
- Conceptual contribution — an ontological reversal. Instead of deriving functional forms from the neuron as a primitive unit and deducing symmetries afterwards, the paper treats symmetry as fundamental and neurons as emergent from group representations. This is formalised in Eqns. 1–3, which taxonomise three "strengths" of symmetry relation over a functional class: algebraic (Eqn. 1), probabilistic (Eqn. 2) and closure (Eqn. 3). The paper states the probabilistic category is not needed for the adaptive networks and is included for completeness.
- Reformulation contribution — isotropic primitives. Using continuous orthogonal equivariance as the symmetry prescription, the paper derives activation functions that commute with every orthogonal matrix R in O(n), i.e. [R, f] = 0 (Eqn. 4). These are "maximally" satisfied by the multivariate forms in Eqn. 5, which are basis-independent and therefore "deindividuated". The author notes O(n) is a supergroup of the permutation group, O(n) ⊃ S_n, making this a minimal natural generalisation.
- Practical contribution — diagonalisation, intrinsic length and adaptive topology. Via SVD of an affine map, A = R Λ Qᵀ, the middle layer of a three-layer perceptron can be reparameterised into diagonal one-to-one connectivity with no functional degradation analytically. This supports neurodegeneration, a thresholded buffer of inactive "scaffold" neurons for neurogenesis, asymptotic 50% parameter sparsification, ICA-based representations for interpretability, and a generalised isotropic-perceptron architecture enabling parallel precomputation of all matrix-vector products with a nested functional class.
- A new tunable parameter — the "intrinsic length". The scalar o is introduced so that the residual bias left behind when a singular value tends to zero can be gauged away, restoring full neurodegeneration symmetry in the Λ_ii → 0 limit (Eqns. 16–17).
Main Findings
- Basis independence removes the individuated neuron. For a generic orthogonal R in O(n)\P, the permutation-group representation is not fixed under conjugation, so constraining a function to that representation selects an absolute frame and makes the classical artificial neuron special. For orthogonal-prescribed isotropic functions, the representation set equals O and no direction is inherently distinguished. The paper describes this as treating neuron individuation as an induced artefact of the chosen group representation rather than an intrinsic property of the vector space.
- Diagonalisation is analytically function-invariant. Applying SVD to the middle affine map of a three-layer perceptron with simultaneous isotropic activations leaves the overall MLP function unchanged, since the isotropic function commutes with the orthogonal matrices. The result is a central layer with one-to-one connectivity (Eqn. 11, Fig. 1). A left-sided partial diagonalisation (Eqn. 13) is stated to be sufficient for adaptive topologies and marginally more computationally efficient, while full diagonalisation may give substantial inference-time benefit through sparsity.
- Asymptotic 50% parameter sparsification of dense networks. Every-second layerwise diagonalisation asymptotically achieves an overall 50% parameter reduction while the model function is identical up to floating-point precision. Eqn. 14 gives the sparsity factor for an odd-depth MLP of 2D+1 affine layers and Eqn. 15 for an even-depth MLP of 2D layers, each of width N, with N² + N parameters per layer. Both are stated to sparsify to 50% asymptotically as N, D → ∞. The author stresses this is not a sparse approximation but an analytical equivalence under the orthogonal basis transform, and that the formulas are upper bounds that joint sparsity optimisation and/or function approximation may improve on.
- Neurogenesis is computationally exact; neurodegeneration is arbitrarily well approximated. Neurogenesis embeds the ℝⁿ activation space into ℝⁿ⁺¹ and keeps a buffer of scaffold neurons; the described procedure is function-invariant. Neurodegeneration produces a well-chosen subspace with reparameterisations that closely approximate function invariance, with the approximation quality set by the threshold 0 < ϑ < 1. The paper states the whole approach is "function-invariant, demonstrated to be computationally identical during neurogenesis, arbitrarily well approximated during neurodegeneration".
- Scaffold neurons are trainable despite zero-valued parameters. The isotropic Jacobian (Eqn. 19) contains a generally non-zero "I term" that the elementwise Jacobian (Eqn. 20) zeroes out. These non-elementwise Jacobians distribute learning gradients globally, so apparent neurons can be trained rapidly even when independent of the model's function. The paper notes only a coordinate singularity exists at the zero vector under a suitable choice of σ̃.
- Orthogonal reparameterisation breaks symmetry in training but not at inference. Unlike traditional permutation reparameterisations, orthogonal transforms show the usual forward-pass function invariance but display variance in their coupling to gradient-descent algorithms after the transform, producing a training-time-only breaking of the reparameterisation symmetry, which is absent at inference time.
- Thresholding and scheduling are required in practice. SVD has an O(m²n) computational scaling dependence for ℝ^{n×m} matrices with m < n, so the paper proposes intermittent scheduling — layerwise decomposition once per epoch — noting this scheduling arises from SVD cost rather than dense gradient computation. A scaffold buffer of size Ξ ∈ ℤ₊ is maintained: exceeding it triggers neurodegeneration of the smallest singular-value neuron, falling short triggers neurogenesis. Because isotropic activations are non-elementwise, different values of Ξ produce tangible training and inference differences. The author warns that constant, predefined growth and pruning should be avoided because small singular values would be repeatedly replaced; a reactive implementation is required.
- Pseudoinverse correction for the linear term. Eqn. 18 gives W′ = W Λ Λ′ᵀ(Λ′ Λ′ᵀ)⁻¹ to preserve the map W′Λ′x⃗ ≈ WΛx⃗ by least squares, with the caveat that a very small threshold can give a poorly conditioned inverse. In the standard case of deleting the smallest singular value, this reduces to a simple column deletion of W′.
- Experimental setup and reported quantities. Experiment One (App. I.1) appends or removes diagonalised neurons from the initial ℝ¹⁰⁰ → ℝ¹⁰⁰ affine layer of a [3072, 100, 100, 10] CIFAR-10-trained classifier over 20 independent repeats, giving 7 × 20 unique networks. The seven columns vary the intrinsic length o and the ψ-vector: (o_none, ψ_decay), (o_constant, ψ_decay), (o_trainable, ψ_decay), (o_trainable, ψ_none), (o_trainable, ψ_constant), (o_trainable, ψ_trainable) and (o_trainable, b⃗_adjusted). The mean initial sum of the layer's singular values, Σ₀, is reported as (0.174 ± 0.001) × 10³ for five configurations and (0.164 ± 0.001) × 10³ for two. The per-configuration ε (mean per-example L₂ distance between classifier logits before and after neuroadaptation, tabulated in units of 10⁻⁴) and Σ₁ (sum after neuroadaptation, tabulated in units of 10³) values are not fully legible in the supplied truncated content; the visible fragment includes the values 75, 2 ± 0 and 0.174 ± 0.001.
- Placeholder activation. The empirical section uses σ(a) = tanh a as a working placeholder "in analogy to standard-tanh", explicitly not assumed to be optimal. The paper notes isotropic-tanh saturates, lim_{α→∞} σ(α) = 1, which alleviates concerns about magnitude growth compensating for small singular values, so zero-tending singular values do suggest negligible functional contribution.
- Scope of the experiments is deliberately narrow. The author states the derivations apply to multilayer perceptrons, so the claims must be assessed against that architecture only, despite the precedence of transformer and convolution state of the art, and that the objective is to verify theoretical claims about function invariance and neuroadaptation rather than overall competitiveness.
Methodology in Plain English
The paper starts from a design principle rather than a formula. Instead of picking a familiar activation function and asking what symmetries it has, it picks a symmetry group — rotations and reflections of an n-dimensional vector space, the orthogonal group O(n) — and asks which nonlinear activation functions are unchanged by it. The answer is a family of functions that act on a vector's length and direction rather than on individual coordinates, so no coordinate has a special status. Because such a function commutes with any orthogonal matrix, an orthogonal change of basis applied to the layer's input can be cancelled by the same change applied on the other side.
That cancellation is what enables the trick at the centre of the paper. Any weight matrix can be written as the product of two orthogonal matrices and a diagonal matrix of singular values (the SVD). Inserting this factorisation into the middle layer of a network, and pushing the orthogonal matrices outward through the isotropic activation functions into the neighbouring affine layers, rewrites the layer as a pure diagonal map: each input channel talks to exactly one output channel, in order of singular value magnitude. Because the orthogonal matrices are absorbed rather than discarded, the network computes the same function.
Once the layer is diagonal, adding and removing capacity becomes simple bookkeeping. To add capacity, append a zero row to the diagonal matrix and a corresponding weight column to the next layer — the new "scaffold" neuron contributes nothing initially, and the new column's initialisation determines whether symmetry breaks spontaneously or is explicitly biased. To remove capacity, delete the row with the smallest singular value and compensate through the new intrinsic-length parameter plus optional bias corrections, so that what is left approximates the original function. A threshold decides which singular values count as negligible, and a buffer size decides when the network should grow or shrink. SVD is run only occasionally because of its cost.
The empirical work is a small-scale check: train a CIFAR-10 classifier, then perform this adapt-and-prune operation on one hidden layer and measure how far the classifier's output logits move, averaged per example. The table varies how the intrinsic length and the additive ψ correction are treated (absent, constant, trainable) to see which settings best preserve function.
Why This Matters
Impact on research. The paper reframes activation-function design as a symmetry-prescription problem, arguing that the space of admissible primitives is much larger than the elementwise forms currently used, and that exploring it can unlock behaviours — here, exact whole-neuron addition and near-exact whole-neuron removal — that permutation-symmetric formulations cannot provide. It also identifies a previously unremarked interaction: orthogonal reparameterisations preserve the forward pass but change the coupling to gradient descent, meaning the transform does not leave training dynamics invariant even though it leaves the function invariant.
Real-world applications (implied by the paper's stated use cases):
- Continual learning systems that append, remove or alter tasks at runtime, which the paper explicitly frames as a broad goal of architecture-based continual learning.
- Memory- and compute-constrained inference, where the 50% asymptotic parameter sparsification is analytically function-preserving and could be applied just before inference time.
- Interpretability and real-time monitoring of networks, via ICA-derived representations (App. E) whose basis is chosen to be meaningful rather than arbitrary, "up to ICA's own concept resolvability".
- Adaptive edge or embedded deployment where a model must restructure in real time in response to changing task demands without retraining from scratch.
Industry relevance. The practical selling points are lower memory footprint at inference for an identical function, the ability to grow and shrink a deployed model reactively rather than by retraining, and a diagonalised, ordered view of a layer that is easier to instrument and monitor. The paper is candid that the empirical work covers only MLPs and does not aim at state-of-the-art competitiveness, so near-term industrial uptake would depend on extensions the author leaves to future work, including a conceptual generalisation to convolution in App. D.
Future Directions
- Extending beyond multilayer perceptrons. The author states the claims must be assessed in reference to MLPs only, and offers only a conceptual generalisation for convolution in App. D, while noting the precedence of transformer and convolution state of the art.
- Joint sparsity optimisation (JSO). Optimising over the multi-layer orthogonal direct-sum Lie group to minimise an objective on parameter count through a continuous proxy — the paper mentions L₁-norm-based approaches among applicable forms — with the Eqn. 14 and 15 sparsity formulas described as upper bounds that JSO or function approximation could improve.
- The optimiser interaction. The training-time-only symmetry breaking under orthogonal transforms is discussed in App. A and described as possibly offering an insightful direction for future study, including its relationship to symmetry-breaking initialisation of scaffold neurons (App. A.3).
- Interpretability and monitoring. Whether ICA-derived representations better represent a network for interpretability and real-time monitoring than the SVD choice, and how far ICA's own concept resolvability limits this, is left open (App. E).
- Practical hyperparameter and testable-claim questions raised by the paper: suitable choices for the threshold ϑ, the scaffold buffer Ξ, and the scaffold initialisation; whether transient correction terms such as ψ⃗ must decay before the next layer is pruned; and whether rarely significant connections tend to produce non-robust maladaptations. The paper also notes that alternatives to singular-value-based neuron-importance assessment, such as gradient magnitude with respect to the loss, are possible and that its own choice is not considered the only option.
Target Audience
This paper is most useful to researchers working on adaptive or growing neural network architectures, continual learning, and biologically inspired plasticity, as well as to those applying group theory and symmetry methods to deep learning. Readers interested in network pruning, sparsification and interpretability will find the diagonalisation result directly relevant. Because it is a theory-first paper with a single MLP-scale empirical check, practitioners looking for immediately deployable gains will need the convolution and non-MLP extensions that the paper only sketches. A working knowledge of linear algebra — SVD, orthogonal matrices, Jacobians — is effectively a prerequisite.
Authors’ abstract
This paper introduces a novel methodology for dynamic networks by leveraging a new symmetry-principled class of primitives, isotropic activation functions. This approach enables real-time neuronal growth and shrinkage of the architectures in response to task demand. This is made possible by network structural changes that are invariant under symmetry reparameterisations, leaving the computation identical under neurogenesis and well approximated under neurodegeneration. This is undertaken by leveraging the isotropic primitives' property of basis independence, resulting in the loss of the individuated neurons implicit in the elementwise functional form. Isotropy thereby allows a freedom in the basis to which layers are decomposed and interpreted as individual artificial neurons. This enables a layer-wise diagonalisation procedure, in which typical interconnected layers, such as dense layers, convolutional kernels, and others, can be reexpressed so that neurons have one-to-one, ordered connectivity within alternating layers. This indicates which one-to-one neuron-to-neuron communications are strongly impactful on overall functionality and which are not. Inconsequential neurons can thus be removed (neurodegeneration), and new inactive scaffold neurons added (neurogenesis) whilst remaining analytically invariant in function. A new tunable model parameter, intrinsic length, is also introduced to ensure this analytical invariance. This approach mathematically equates connectivity pruning with neurodegeneration. The diagonalisation also offers new possibilities for mechanistic interpretability into isotropic networks, and it is demonstrated that isotropic dense networks can asymptotically reach a sparsity factor of 50% whilst retaining exact network functionality. Finally, the construction is generalised, demonstrating a nested functional class for this form of isotropic primitive architectures.