Skip to content
AI.info

Research

Can Large Language Models Generalize Procedures Across Representations?

Overview Research area: Natural Language Processing / LLM post-training, specifically cross-representation generalization of procedural knowledge (natural language, code, graphs). Technical level: Int

Can Large Language Models Generalize Procedures Across Representations?
arXiv
2602.03542
Published
2026-02-03
Authors
Fangru Lin, Valentin Hofmann, Xingchen Wan, Weixing Wang, Zifeng Ding, Anthony G. Cohn, Janet B. Pierrehumbert

AI summary

Overview

Research area: Natural Language Processing / LLM post-training, specifically cross-representation generalization of procedural knowledge (natural language, code, graphs).

Technical level: Intermediate. The paper uses standard post-training methods (SFT, distillation, STaR, GRPO) and a curriculum-learning setup; the analysis section draws on structure-mapping theory and graph kernels, which require some background.

Scope in one sentence: The paper asks whether LLMs trained on symbolic representations (code, graphs) can transfer the underlying procedures to isomorphic natural-language tasks, and proposes a two-stage RL curriculum that makes such transfer work.

What This Paper Is About

LLMs are trained heavily on symbolic data such as code and graphs, but real users typically state tasks in natural language. The paper tests whether a procedure learned in one representation carries over to another when the underlying task is exactly the same — for example, estimating the shortest completion time for a set of interdependent planning steps, which is equivalent to finding the critical path in a Directed Acyclic Graph (DAG) regardless of how it is written down. The goal is to isolate procedure learning from surface-form learning, then find a training recipe that actually produces cross-representation generalization.

Key Contributions

  1. A controlled isomorphic testbed. The authors build natural language (NL), Graph, and Code versions of the same asynchronous planning task from AsyncHow data, so surface form varies while the underlying algorithm stays fixed. Graph is an adjacency-list dictionary with dummy START and END nodes plus a separate time dictionary; Code is a Python DAG longest-path function with weighted edges and times converted to minutes.

  2. A systematic negative result. They train Qwen-2.5-Instruct (1.5/3/7B), Llama-3.2-1/3B-Instruct, Llama-3.1-8B-Instruct, and Olmo-2-1/7B-Instruct with vanilla SFT, distillation from DeepSeek-R1-Distill-Qwen-32B, STaR, and GRPO, training on one representation and testing on all three. Training on symbolic data alone does not reliably transfer procedures to natural language across any method, model family, or scale.

  3. A two-stage RL curriculum. Training first on symbolic data (Graph) and then on natural language with GRPO substantially improves NL performance at a fixed token budget, outperforming NL-only training and even GPT-4o-mini, and matching zero-shot GPT-4o on the planning task.

  4. An analogy-based explanation. Using structure-mapping theory, the authors show that successful cross-representation generalization correlates better with the analogical strength of the most similar learned training item than with the frequency of moderately similar items — and that the curriculum amplifies this analogical behavior.

Main Findings

  • Symbolic training does not transfer to natural language. Across all LLMs and post-training methods tested, training on a single representation fails to reliably generalize across representations. Models show high within-representation performance but performance collapses under representation shift, even though the underlying procedures are identical. Even when transfer is statistically significant (e.g., Qwen-2.5-Instruct-1.5B trained on Graph with STaR), out-of-representation performance remains markedly lower than within-representation performance.

  • Training representation matters. NL training gives the strongest NL test performance, but small models still struggle within-representation, sitting around 0.5 accuracy at the 1.5B and 3B scales. Graph and Code training achieve high within-representation accuracy but transfer poorly to NL.

  • RL is best in-distribution but not under shift. GRPO has the strongest within-representation performance across methods, but its relative advantage diminishes under representation shift despite significantly more training compute. Among SFT methods, vanilla SFT is worst overall; distillation substantially improves within-representation performance and approaches GRPO with less budget but sometimes exacerbates cross-representation degradation; STaR generalizes better than vanilla SFT but remains weaker than distillation.

  • Scaling does not fix it. Larger models show the same qualitative pattern, with a more noticeable performance drop in unseen representations that is likely due to their stronger in-representation performance rather than weaker transfer.

  • The curriculum works. Training a Qwen2.5-1.5B model with GRPO on Graph for 40 steps (20 episodes), then NL for 40 steps, gives 0.873 train accuracy on NL and 0.782 test accuracy on NL, versus 0.811 and 0.698 for an identical model trained on NL only for 80 steps. The curriculum model also reaches 0.573 on NL-AAVE versus 0.507 for NL-only.

  • Efficient scaling benefit. The 80-step 1.5B curriculum model outperforms a 3B model trained on NL alone for 40 steps (0.471 NL test accuracy) and outperforms its own 7B variant (0.698 NL test accuracy), which requires approximately 2.3 times as much training budget. It outperforms zero-shot GPT-4o-mini (0.440) and matches zero-shot GPT-4o (0.782) on the asynchronous planning task, though GPT-4o leads on NL-AAVE (0.724 versus 0.573).

  • Order and method both matter. Reversing the curriculum to NL then Graph performs markedly worse than NL-only (0.431 versus 0.698 on NL). Replacing stage 2 GRPO with NL distillation gives 0.462. The curriculum does not work under SFT either (0.236 for NL-only at 2 epochs versus 0.244 at 4 epochs versus 0.249 for Graph 2 epochs then NL 2 epochs), and parameter-efficient SFT with rank 8/16 LoRA shows no meaningful improvement.

  • Interleaving is worse than sequencing. Interleaved Graph+NL training for 40 gradient steps reaches 0.382 on NL, versus 0.782 for the Graph then NL curriculum.

  • A strong first stage is essential. Code alone trains slowly (0.338 accuracy after 40 steps), and Code (40 steps) then NL (40 steps) is worse than NL-only (40 steps) (0.382 versus 0.538). Graph+Code then NL is worse than Graph then NL (0.533 versus 0.782) because the first phase is too weak (0.522). Graph (40 steps) then Code (40 steps) reaches 0.787 when tested in Code, versus 0.373 for Code-only (80 steps).

  • Robustness to dialect. On NL-AAVE, the curriculum (0.573) outperforms NL-only (0.507), the reversed NL then Graph curriculum (0.169), GPT-4o-mini (0.289), and a 3B Qwen model trained on NL only (0.400), despite having no explicit exposure to the dialect.

  • Generalizes to math and physics. On MATH levels 4 and 5 and SciBench physics questions, Code then NL with GRPO generalizes better across domains than NL-only, even when in-domain performance is comparable or slightly lower. Math-trained curriculum reaches 0.435 on Math and 0.230 on Physics versus 0.385 and 0.135 for NL-only; physics-trained curriculum reaches 0.550 on Physics and 0.325 on Math versus 0.555 and 0.315 for NL-only. The curriculum also produces consistent gains on Olmo-2-7B-Instruct.

  • Success looks like analogy, not frequency. Analogy-based correlations exceed frequency-based ones in every setting: train NL / test NL (0.242 versus 0.176), train NL / test Graph (0.148 versus 0.124), train Graph then NL / test NL (0.265 versus 0.245), and train Graph then NL / test Graph (0.297 versus 0.273). When trained on Graph and tested on NL, neither hypothesis explains results, consistent with the lack of visible transfer. NL test performance after curriculum training correlates more with Graph within-representation training than with NL within-representation training (0.526 versus 0.353, both p < 0.001).

  • Qualitative behavior differs by training regime. A Graph-only model fails to recognize the underlying graph structure in NL and defaults to summing all time constraints. An NL-trained model understands it must search for the critical path. The curriculum model iterates across multiple candidate paths and compares them, rather than committing to a single path — though it still makes errors such as wrong time unit conversion, suggesting a ceiling set by base model capability.

Methodology in Plain English

The authors take a planning task that can be written three ways without changing what has to be computed. In natural language, a person describes making a dish, lists steps with durations, and states which steps must come before others; the answer is the shortest possible completion time assuming infinite resources. That same problem is a DAG where nodes are steps and edges are dependencies, so the answer is the length of the longest directed path. The authors then render it as an adjacency list dictionary (Graph) and as a Python longest-path function with numeric edge weights (Code).

Because all three versions share the same underlying structure, any performance difference under a representation switch must come from generalization of the procedure, not from spurious surface cues. The authors train base models on exactly one representation and test on all three, using training and test splits of 1,364 and 225 data points after deduplication with stratified sampling by complexity. They judge significant transfer with McNemar's tests against untuned baselines.

To fix the transfer failure, they split learning into two stages: symbolic induction first, so the model learns the abstract procedure, then natural language adaptation, so the procedure becomes usable in NL. They use GRPO with verifiable outcome rewards and keep the total training token budget fixed across conditions to make comparisons fair.

For analysis, they quantify how similar a training instance is to a test instance using an analogical strength measure from structure-mapping theory, weighting binary relational similarity more than unary item similarity with a discount factor of 0.4, measuring unary similarity by multi-set histogram-Jaccard over node time durations and binary similarity by a Weisfeiler–Lehman subtree kernel with 3 iterations. They then correlate test accuracy with either the count of moderately similar learned items (frequency hypothesis) or the similarity of the k-th most similar learned item (analogy hypothesis), sweeping k from 1 to 10 and p from 0.1 to 0.9 and reporting Pearson's correlation.

Why This Matters

Impact on research. The paper provides a clean, controlled counterexample to the assumption that symbolic training such as code or graph data improves natural-language reasoning. It shows that high within-representation accuracy does not imply transferable procedural knowledge, that RL's in-distribution advantage does not survive a representation shift, and that model scaling does not fix the problem. It also connects LLM generalization to cognitive-science accounts of generative analogy, offering a measurable predictor of when transfer will succeed.

Real-world applications:

  • Planning and scheduling assistants that must reason about task dependencies stated by users in ordinary language.
  • Models trained predominantly on code being deployed to answer natural-language requests, where the paper's results caution against assuming inherited capability.
  • Robustness work for dialectal and non-standard language, since the curriculum-trained model improved on NL-AAVE without any dialect training data, though a notable gap to NL remains.
  • Education and tutoring systems that need to move learners, or models, between symbolic formalisms and verbal explanations of the same procedure.

Industry relevance. The curriculum reaches GPT-4o-level NL planning performance with a 1.5B model, which matters for cost and deployment size. The finding that curriculum-trained models beat larger models trained conventionally at the same budget, and that a 7B model needs roughly 2.3 times the budget, is directly relevant to training-cost planning. The null results for SFT, LoRA (rank 8/16), and interleaved data mixtures also warn teams that common recipe choices will not buy cross-representation generalization.

Future Directions

  • Scaling representations at training time. The impact statement explicitly notes it is an open question whether training with more representations yields an observable advantage.
  • Closing the NL versus NL-AAVE gap. The curriculum improves dialect performance but the gap remains notable, which the authors frame as a technological fairness issue warranting better solutions.
  • Reducing the training requirement. The authors contrast LLM behavior with human analogical reasoning, which generalizes across representations with minimal or zero exposure; narrowing that gap is an explicit open problem.
  • Explaining representation asymmetry. Graph is a strong first stage while Code is not (0.338 after 40 steps), and Graph then Code beats Code-only when tested in Code. Understanding why some symbolic representations provide a better inductive foundation would let the curriculum be applied more reliably across tasks.

Target Audience

Researchers and engineers working on LLM post-training, reinforcement learning for reasoning, and symbolic-to-natural-language transfer. It is also relevant to cognitive scientists interested in analogy and structure mapping as computational accounts of generalization, and to practitioners choosing between SFT, distillation, STaR, and GRPO for training data with mixed representations.

Authors’ abstract

Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are often specified in natural language. To what extent can LLMs generalize across these representations? Here, we approach this question by studying isomorphic tasks involving procedures represented in code, graphs, and natural language (e.g., scheduling steps in planning). We find that training LLMs with popular post-training methods on graphs or code data alone does not reliably generalize to corresponding natural language tasks, while training solely on natural language can lead to inefficient performance gains. To address this gap, we propose a two-stage reinforcement learning curriculum that first trains on symbolic, then natural language data. The curriculum substantially improves model performance across model families and tasks. Remarkably, a 1.5B Qwen model trained by our method can closely match zero-shot GPT-4o in naturalistic planning. Finally, our analysis suggests that successful cross-representation generalization can be interpreted as a form of generative analogy, which our curriculum effectively encourages. The dataset and code used in this paper can be found \href{https://github.com/fangru-lin/procedure_generalization_llm}{here}.

Read the original paper