Research
Iterative Training of Physics-Informed Neural Networks with Fourier-enhanced Features
Overview Research area: Physics-informed machine learning — specifically, mitigating spectral bias in Physics-Informed Neural Networks (PINNs) for solving partial differential equations (PDEs). Techni

- arXiv
- 2510.19399
- Published
- 2025-10-22
- Authors
- Yulun Wu, Miguel Aguiar, Karl H. Johansson, Matthieu Barreau
AI summary
Overview
Research area: Physics-informed machine learning — specifically, mitigating spectral bias in Physics-Informed Neural Networks (PINNs) for solving partial differential equations (PDEs).
Technical level: Advanced. The paper combines a bi-level optimization framework, convergence proofs for a strongly convex quadratic lower-level problem, Lipschitz continuity of the solution map, and universal approximation arguments, alongside empirical benchmarks.
Scope: The paper proposes IFeF-PINN, an iterative two-stage training algorithm that extends a PINN's latent basis with random Fourier features and solves the resulting coefficient regression exactly for linear PDEs, then evaluates it on four benchmark PDEs against four baselines.
What This Paper Is About
PINNs approximate PDE solutions with neural networks, but they suffer from spectral bias: they learn low-frequency components first and often fail to capture oscillatory, high-frequency behavior. The paper reframes PINN training as a two-stage problem — building a feature basis, then regressing onto its coefficients — and enriches that basis with random Fourier features (RFF) so the network can represent high-frequency content. The goal is a method that is both more accurate than existing PINN variants on high-frequency and multi-scale PDEs and theoretically grounded for linear PDEs.
Key Contributions
-
A flexible architectural building block. An RFF mapping applied to the last hidden layer features of a PINN, which upgrades the implicit dot-product kernel on those features to a stationary kernel and injects high-frequency components without adding trainable parameters. The authors demonstrate its universal approximation capabilities.
-
A two-stage iterative training algorithm with convergence guarantees. Basis generation (upper level) is decoupled from linear regression on the basis coefficients (lower level). For linear PDEs the lower-level problem is a strongly convex quadratic program with a closed-form solution, and the paper proves that the iterative scheme converges to a stationary point.
-
Extensive numerical validation. Benchmarks on the 2D Helmholtz equation (low and high frequency), the 1D convection equation (low and high frequency), the 1D convection-diffusion equation, and the viscous Burgers' equation, against Vanilla PINNs, NTK, PINNsformer, and Physics-Informed Gaussians (PIG) as baselines.
-
Generalization across training strategies. IFeF-PINN is shown to compose with advanced PINN techniques, demonstrated with a primal-dual weight-balancing variant (IFeF-PD).
Main Findings
-
Best low-frequency accuracy. On the low-frequency 2D Helmholtz problem, IFeF-PD attains the best relative L2 error of 3.5 × 10⁻⁵; on the low-frequency 1D convection problem, IFeF attains 4.3 × 10⁻⁵. Across the three low-frequency benchmarks (Helmholtz, convection, viscous Burgers), the proposed method achieves the lowest median errors with reduced variability.
-
Decisive gains on high-frequency problems. For the high-frequency Helmholtz equation with a₁ = a₂ = 100, Vanilla, PINNsformer, and NTK all failed to converge (marked "-"), and PIG reached 1.6884 (0.2775); IFeF reached 0.0156 (0.0055) and IFeF-PD reached 0.0092 (0.0031). For the convection equation with β = 200, error dropped from 0.9024 (0.0239) for Vanilla, 1.2278 (0.2010) for PINNsformer, 0.8685 (0.0318) for NTK, and 1.0009 (0.0003) for PIG, to 0.0027 (0.0010) for IFeF and 0.0025 (0.0005) for IFeF-PD.
-
Multi-scale case. On the convection-diffusion equation (k_low = 4π, k_high = 60π), baselines clustered near 0.05 (Vanilla 0.0501 (0.0030), PINNsformer 0.0525 (0.0001), NTK 0.0526 (0.0001), PIG 0.0560 (0.0010)), while IFeF reached 0.0009 (0.0003) and IFeF-PD 0.0010 (0.0002). The authors state only IFeF-PINN learned both the low- and high-frequency components.
-
The RFF extension is the essential ingredient. An ablation that removed the RFF basis extension but kept the iterative two-step optimization produced results only slightly better than Vanilla PINN (1.4923 × 10⁻² relative L2 error) on the low-frequency convection problem, and failed to converge entirely on the high-frequency Helmholtz and convection equations.
-
Two-stage training beats end-to-end training. Jointly optimizing ω and θ (end-to-end) gave 0.0088 (0.0006) on Helmholtz (a₁ = 1, a₂ = 4) and 0.0049 (0.0009) on viscous Burgers (ν = 0.01/π), and failed to converge on Helmholtz (a₁ = a₂ = 100). IFeF gave 0.0003 (0.0003), 0.0024 (0.0011), and 0.0156 (0.0055) respectively; IFeF-PD gave 0.00005 (0.00002), 0.0033 (0.0004), and 0.0092 (0.0031). Improvements were large for linear PDEs and modest for nonlinear PDEs where the lower level becomes non-convex.
-
Spectrum analysis confirms the mechanism. Using an initial condition built from a superposition of ten sinusoids of different frequencies with unit amplitude, the network's discrete Fourier transform magnitudes at frequency k_i were compared. Vanilla PINNs struggled on high-frequency components; extending the basis with RFF improved high-frequency fitting even without the subsequent bi-level training, and increasing the number of random features further improved high-frequency approximation. Magnitudes were averaged over five independent runs.
-
Hyperparameter sensitivity. On low-frequency Helmholtz (a₁ = 1, a₂ = 4) with σ = 1, error was 5.5 × 10⁻⁴ at D = 100, 2.1 × 10⁻⁴ at D = 400, 3.2 × 10⁻⁴ at D = 800, 5.7 × 10⁻⁴ at D = 1200, and 4.5 × 10⁻⁴ at D = 3000. With D = 800, error across σ ∈ {2, 1, 0.5, 0.2, 0.1} was 4.0 × 10⁻⁴, 3.2 × 10⁻⁴, 5.5 × 10⁻⁴, 3.3 × 10⁻⁴, and 1.5 × 10⁻³. For high-frequency Helmholtz (a₁ = a₂ = 100) with σ = 1, error fell from 7.11 × 10⁻² at D = 800 to 1.56 × 10⁻² at D = 2400, rising again to 2.22 × 10⁻² at D = 3000; with D = 2400, error across σ ∈ {20, 10, 5, 1, 0.2} was 4.6 × 10⁻³, 3.0 × 10⁻³, 5.7 × 10⁻³, 1.56 × 10⁻², and 1.05 × 10⁻¹. Low-frequency problems were robust to both hyperparameters; high-frequency problems were sensitive, especially to σ.
-
Ablation caveats stated by the authors. Too few Fourier features reduce expressivity, while excessive features cause overfitting and may break the rank condition required for the positive-definite Q matrix. Larger σ values are essential for high-frequency problems.
Methodology in Plain English
The starting point is the observation that standard PINN training conflates two jobs inside one nonconvex objective: the hidden layers learn a nonlinear feature basis, and the final linear layer finds coefficients that project the solution onto that basis. The authors split these apart.
Stage one (basis generation). Train a standard PINN as a warm start, producing hidden-layer features h_ω that serve as a nominal basis. The authors note this warm start is necessary for homogeneous PDEs to prevent the network from converging to u ≡ 0, since standard initialization yields near-zero outputs that trivially minimize the lower-level problem; the source term in non-homogeneous PDEs prevents this issue.
Basis extension. Apply a random Fourier feature mapping γ_D to those hidden features, multiplying them by a fixed random matrix B_D whose entries are drawn i.i.d. from N(0, σ²), and taking sines and cosines. This turns the features into an extended basis ψ_D. It upgrades the implicit dot-product kernel on h_ω into a stationary kernel in the adaptive feature space, and it increases the number of basis functions independently of the network's width — with no new trainable parameters.
Stage two (lower-level regression). Approximate the solution as a linear combination of the extended basis, u_{ω,θ}(x) = ψ_D(x)ᵀθ. Because the differential and boundary operators are linear, the loss becomes quadratic in θ: ½θᵀQ(ω)θ + c(ω)ᵀθ + b. When λ > 0 and a rank condition holds, Q is positive definite and the unique optimum is θ* = −Q⁻¹c. This exact solve is executed every epoch, followed by one gradient-descent step on ω, alternating until convergence.
Nonlinear PDEs. For nonlinear operators, the lower-level problem stops being convex and has no closed-form solution. The algorithm is modified to find an approximate local minimizer by gradient descent, refreshed only every N_lower epochs, with the local minimizer retaining Lipschitz continuity and differentiability when the Second-Order Sufficient Condition holds.
Theory. The convergence argument proceeds in two steps: first, showing the optimal lower-level solution map θ*(ω) is locally Lipschitz continuous in ω (given locally Lipschitz Q and c and a uniform lower bound μ_Q > 0 on the smallest eigenvalue of Q); second, showing the hypergradient — derived via the Implicit Function Theorem — is L-smooth, so constant-step-size gradient descent with η ∈ (0, 2/L) converges to a stationary point. Separately, the paper proves that the composite RFF function space has projection error no greater than the original feature space as D → ∞, yielding a universal approximation corollary given enough neurons and RFF features.
Experimental protocol. Relative L2 error ‖u_pred − u_real‖₂ / ‖u_real‖₂ is measured after convergence. Each method is run five times with independent random seeds, and the best predictions for each approach are reported. λ = 0.01 for Vanilla PINNs. Uniform sampling follows Zhao et al. (2024) for the low-frequency Helmholtz and low-frequency convection problems; the viscous Burgers setup follows Raissi et al. (2019); Latin hypercube sampling is used for the high-frequency Helmholtz equation. All models are implemented in PyTorch and trained on a single NVIDIA GeForce RTX 4090 GPU. Code is available at https://github.com/CyberAltrumi/IFeF-PINN.
Why This Matters
Impact on research. The work reframes spectral bias not as a weighting or sampling problem but as a basis-representation problem, then supplies convergence guarantees that single-level PINN training generally lacks. It positions itself against weight-balancing methods (weak convergence guarantees, still single-level), resampling methods (do not explicitly target spectral bias), and curriculum/architecture methods (complex, slow, no dedicated optimization algorithms). The authors note the framework is complementary to weight balancing and can incorporate it, as demonstrated by IFeF-PD.
Real-world applications (drawn from the phenomena the paper cites as requiring high-frequency capture):
- Wave propagation modeling.
- Turbulence simulation.
- Quantum dynamics.
- General PDE solution as a grid-free surrogate, benefiting from adaptability to complex geometries and high-dimensional scalability noted for the PINN paradigm.
Industry relevance. Many engineering and scientific workflows rely on solving PDEs where oscillatory or multi-scale behavior dominates, and where mesh-based spectral, multiscale, or oscillatory quadrature methods require problem-specific adaptation or become prohibitively costly in complex or high-dimensional settings. A grid-free neural surrogate that reliably resolves high frequencies reduces the need for bespoke numerical schemes. The measured gains are substantial where baselines fail: on high-frequency Helmholtz, three of four baselines did not converge at all while IFeF-PD reached 0.0092. The stated trade-offs are a high memory footprint and unsuitability for resampling-based approaches.
Future Directions
-
Principled bi-level optimization for nonlinear PDEs. The authors identify this as the main open problem: the nonlinear lower-level problem is nonconvex, precluding a one-step solve, requiring iterative two-stage gradient descent updates that can stall in local minima.
-
Reducing memory footprint. The paper lists a high memory footprint as an explicit limitation of the method, and notes that it is not adapted to resampling strategies.
-
Balancing D and σ systematically. The hyperparameter ablation shows conflicting pressures — too few features reduce expressivity, too many risk overfitting and can break the rank condition, and performance is non-monotonic in D for the high-frequency Helmholtz case (best at 2400, worse at 3000). Low-frequency problems are robust to both hyperparameters while high-frequency problems are sensitive, especially to σ.
-
Composition with other PINN advances. The paper treats weight balancing, resampling, and curriculum/architecture methods as complementary and demonstrates integration with primal-dual weight balancing; further combinations with these single-level techniques remain open.
Target Audience
Researchers and graduate students working on physics-informed machine learning, scientific computing with neural surrogates, and spectral bias in neural networks. It is most useful to readers comfortable with PDE-constrained optimization, bi-level programming, and kernel methods (random Fourier features, Bochner's theorem, reproducing kernel Hilbert spaces), since the theoretical sections rely on strong convexity, Lipschitz continuity, the Implicit Function Theorem, and universal approximation arguments. Practitioners seeking a drop-in method for high-frequency PDE surrogates will find the algorithm descriptions and the public code repository directly actionable, though they should note the memory-footprint limitation and the reduced guarantees for nonlinear PDEs.
Authors’ abstract
Spectral bias, the tendency of neural networks to learn low-frequency features first, is a well-known issue with many training algorithms for physics-informed neural networks (PINNs). To overcome this issue, we propose IFeF-PINN, an algorithm for iterative training of PINNs with Fourier-enhanced features. The key idea is to enrich the latent space using high-frequency components through Random Fourier Features. This creates a two-stage training problem: (i) estimate a basis in the feature space, and (ii) perform regression to determine the coefficients of the enhanced basis functions. For an underlying linear model, it is shown that the latter problem is convex, and we prove that the iterative training scheme converges. Furthermore, we empirically establish that Random Fourier Features enhance the expressive capacity of the network, enabling accurate approximation of high-frequency PDEs. Through extensive numerical evaluation on classical benchmark problems, the superior performance of our method over state-of-the-art algorithms is shown, and the improved approximation across the frequency domain is illustrated.