Research
Prompting Neural-Guided Equation Discovery Based on Residuals
Overview Research area: Machine learning / symbolic regression — specifically equation discovery systems (EDSs), which predict a mathematical equation that describes a data set. Technical level: Inter

- arXiv
- 2511.05586
- Published
- 2025-11-05
- Authors
- Jannis Brugger, Viktor Pfanschilling, David Richter, Mira Mezini, Stefan Kramer
AI summary
Overview
Research area: Machine learning / symbolic regression — specifically equation discovery systems (EDSs), which predict a mathematical equation that describes a data set.
Technical level: Intermediate. The core idea is intuitive, but following the method requires basic familiarity with syntax trees, residuals, and symbolic regression benchmarks.
Scope: The paper introduces and evaluates RED (Residuals for Equation Discovery), a post-processing method that uses a syntax tree of an initially predicted equation to compute residuals for its subequations, re-prompts an EDS with those residuals as a new target variable, and replaces a subequation when the replacement lowers validation error.
The preprint is arXiv:2511.05586v1 [cs.LG], dated 05 Nov 2025, licensed CC BY-SA 4.0. The paper states it has not undergone peer review, and that the Version of Record is published in Lecture Notes in Computer Science, volume 16090 (doi.org/10.1007/978-3-032-05461-6_7).
What This Paper Is About
Neural-guided equation discovery systems take a data set as a prompt and predict an equation without extensive search, but if the predicted equation does not meet the user's expectations, there are few options for getting other suggestions without intensive work with the system. RED fills this gap by computing, for each subequation inside the initial equation, what that subequation should have produced for every data point so the whole formula predicts correctly — these values become a new target variable and a new prompt. If the EDS then proposes a subequation with lower validation error than the old one, RED swaps it into the original equation.
Key Contributions
-
A node-based residual calculation scheme. By parsing an equation into a syntax tree with an added Y-node as parent of the root, the authors define the mathematical behavior of each operator node depending on which adjacent node calls it, so residuals can be computed as fast as evaluating the tree, and new mathematical operations can be added as new node types.
-
The RED algorithm. RED iterates over nodes in a breadth-first ordering (left to right within a depth, then deeper), computes the residual for a node, asks the EDS to predict a new subtree, substitutes it into the tree, and accepts it only if the validation error decreases. The child of the Y-node is excluded because its residual is equivalent to the original problem, and the node just updated plus its direct parents are removed from the residual list.
-
A model-agnostic design. RED is stated to be usable with any equation discovery system, fast to calculate, and easy to extend for new mathematical operations.
-
An empirical evaluation on 53 Feynman data sets comparing RED against six other post-processing methods across five EDSs, plus a sensitivity analysis of iteration count, noise, and data set size.
Main Findings
-
RED improves every tested system. On 53 Feynman data sets, RED helped all tested neural-guided EDSs and all tested classical genetic programming systems.
-
Win ratios for NeSymReS. In one-vs-one comparisons over all equations each method proposed, Seeded GPLearn won more often than it lost against all methods, but the difference compared to RED was only minimal; the paper states that RED's equations are of higher quality, giving a higher win ratio.
-
SymbolicGPT benefits least. RED's improvement for SymbolicGPT was not as drastic as for the other EDSs, which the paper attributes to RED's dependence on the initial equation — SymbolicGPT's initial equations had the highest error, and a poor initial equation can make the residual problem harder than the original.
-
Median error table (selected figures). For NeSymReS, Classic achieved MSE Q2 = 0.46 and Q3 = 10.24 with 4 operators (median) and a median runtime per equation of 3.27 sec, while RED achieved MSE Q2 = 0.30 and Q3 = 5.46 with 8 operators (median) and 22.68 sec. For PySR, Classic gave MSE Q2 = 0.05, Q3 = 0.14 and RED gave 0.02, 0.12. For GPLearn, Classic gave 0.05, 0.99 and RED gave 0.02, 0.33.
-
RED lengthens equations. A drawback the paper highlights is that the average length of equations produced with RED increases significantly; the RED row for NeSymReS has a median of 8 operators versus 4 for Classic.
-
Iteration count. With the maximum set to 100, the relative test error for NeSymReS dropped sharply within a few iterations and reached a minimum, while the average number of operators in the equation increased with the number of iterations. The authors selected a maximum of 10 iterations to balance runtime, overfitting, and accuracy.
-
Noise behavior. Adding relative uniform noise to each data set value, the training error for Permute increased linearly with noise, while RED's training error remained almost constant, leading to similar test-data values for the two approaches. From a certain noise level on, RED tends to overfit. Increasing noise decreased the median number of operations for RED but increased it for Permute.
-
Data set size. Running data sets of size 10, 20, 50, 100, 200, 300 and 500, Permute's train and test error decreased with more data points. Combined with RED, the MSE remained at a constant low level regardless of the number of samples. The paper notes NeSymReS's inability to fit few data points was surprising and probably because the neural model was not trained for such small data sets.
-
Other post-processing methods. Seeded GPLearn improved all examined EDSs. Hyper Parameter Grid, Fitting, and Permute showed promising results especially with the genetic algorithms, but with NeSymReS gave only medium improvements. CVGP could only keep up regarding the median value for GP and was otherwise significantly worse. E2E was tested with only one seed and then excluded, because it had to be restarted multiple times and had the longest suggested equations and running time.
-
Completion and failure modes. Only PySR provided a solution for all 53 equations and NeSymReS for almost all (52 completed for Classic); both were also the fastest. Reported failures include token-based neural approaches predicting semantically unsound expressions such as sinI, over- or underflows, divisions by zero, undefined exponentiation in R, negative logarithms, and equations without an operator (which prevent applying residuals).
-
Total running times. Aggregate running times over 3 seeds were 3735 sec for NeSymReS, 4512 sec for SymbolicGPT, 1823 sec for PySR, and 12717 sec for GPLearn.
Methodology in Plain English
The authors start from an equation already produced by an EDS and turn it into a syntax tree: inner nodes are operators and functions, leaves are variables and constants, and a special Y-node is added above the root to hold the data set's output column. Because each node corresponds to a subequation, the tree can be rewritten as an equation system with one operation per equation. Rearranging that system toward any chosen node tells you what that node's subequation must have produced for each data point so the whole formula still equals y — this is the residual, and computing it only requires that all operators on the path from the node to the root are invertible. Operators that are not bijective, such as sine, cannot be inverted, so residuals cannot be computed for their children.
RED walks the tree in a fixed order, computes a residual for one node, and treats the residual as a new target column for the EDS. The EDS returns a candidate subtree, which is swapped into the tree; the modified equation is then scored on a held-out validation set, and the swap is kept only if the error dropped. When a swap is kept, the process restarts on the updated tree, skipping the updated node and its direct parents since their residuals are unchanged or equivalent to previously considered problems. The loop stops when the error threshold is met, the iteration limit is reached, or no residual nodes remain.
The test protocol used 300 rows sampled per data set, split 60% training, 20% validation, 20% test, with MSE as the test metric; a data set counts as unsolved if Classic's MSE exceeds 0.001, and only unsolved data sets are used to test post-processing. Experiments used 3 seeds, except E2E with one seed.
Why This Matters
Impact on research. RED is a general, cheap-to-compute wrapper that can be attached to any equation discovery system rather than replacing it, and the paper reports gains for both neural-guided and genetic programming approaches. It also contributes a conceptual framing — an equation as a system of subequations from which one can factor out the already-solved parts — that connects to prior decomposition ideas such as AI Feynman and AI Feynman 2.0, additive regression, BACON, and CVGP, and to prompt-refinement ideas from test-time augmentation and LLM prompt engineering.
Real-world applications. The paper does not report specific deployed applications, but the setting it studies — recovering interpretable equations from data — is common in:
- Scientific modeling where the goal is a readable formula rather than a black-box predictor.
- Experimental data analysis where a first equation is close but not satisfying and the user wants targeted alternatives.
- Automation pipelines that need repeatable equation suggestions rather than manual re-runs of a discovery tool.
- Settings where small data sets or noisy measurements make a one-shot prediction unreliable.
Industry relevance. Because RED requires no retraining and works on top of existing EDSs, it is a low-friction addition to tooling built around systems such as PySR, GPLearn, or pretrained neural equation predictors. The paper's own caveats matter here: RED increases equation length, adds runtime per equation (for example 22.68 sec median for RED versus 3.27 sec for Classic with NeSymReS), and needs a validation set to decide whether to accept a replacement.
Future Directions
-
Control equation growth. Using equation length, not only validation error, as an acceptance criterion for new subequations to attenuate RED's tendency toward long equations.
-
Parallel and dynamically ranked search. A parallel search combined with a dynamic ranking is proposed to achieve improvements within the same running time.
-
Better integration with genetic programming. The paper notes that the advantage of decomposing the original problem is less evident for random-search genetic approaches, and suggests using RED as an operator inside genetic algorithms, plus studying the mutual influence of genetic EDSs and RED.
-
Dependence on the initial equation. RED can fail when the initial equation does not enable disentanglement and the residual problem becomes harder than the original; the authors expect that better neural-guided EDSs would make RED even more effective.
Target Audience
Researchers and practitioners in symbolic regression and equation discovery who already use EDSs and want a post-processing improvement without changing the underlying model; machine learning engineers building automated modeling pipelines; and scientists who need interpretable equations from data. Readers interested in the internals of tree-based residual computation, or in the empirical comparison of post-processing strategies on the Feynman benchmark, will benefit most.
Authors’ abstract
Neural-guided equation discovery systems use a data set as prompt and predict an equation that describes the data set without extensive search. However, if the equation does not meet the user's expectations, there are few options for getting other equation suggestions without intensive work with the system. To fill this gap, we propose Residuals for Equation Discovery (RED), a post-processing method that improves a given equation in a targeted manner, based on its residuals. By parsing the initial equation to a syntax tree, we can use node-based calculation rules to compute the residual for each subequation of the initial equation. It is then possible to use this residual as new target variable in the original data set and generate a new prompt. If, with the new prompt, the equation discovery system suggests a subequation better than the old subequation on a validation set, we replace the latter by the former. RED is usable with any equation discovery system, is fast to calculate, and is easy to extend for new mathematical operations. In experiments on 53 equations from the Feynman benchmark, we show that it not only helps to improve all tested neural-guided systems, but also all tested classical genetic programming systems.