Research
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
Overview Research area: Natural language processing; specifically the pre-training of large language models and the "pre-pretraining" (PPT) warm-up phase on synthetic, non-natural-language data. Techn

- arXiv
- 2609.39827
- Published
- 2026-09-30
- Authors
- Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal Štefánik, Aline Villavicencio, Nikolaos Aletras
AI summary
Overview
- Research area: Natural language processing; specifically the pre-training of large language models and the "pre-pretraining" (PPT) warm-up phase on synthetic, non-natural-language data.
- Technical level: Advanced.
- Scope: A large-scale empirical study testing whether synthetic pre-pretraining retains its benefits at 500M–7B parameters, across four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, and whether the claimed explanation (a "grammatical prior") holds.
What This Paper Is About
Recent work claims that warming up a language model on synthetic non-natural sequences such as the bracket-matching language k-Shuffle Dyck improves token efficiency in later pre-training, and attributes this to a grammatical prior — a structural inductive bias that transfers to natural language grammar. That evidence, however, comes from models of at most 1B parameters, pre-training budgets below 2B tokens, and predominantly web-text corpora. This paper asks whether the gains survive larger models, longer training, and realistic multi-domain pre-training mixtures (web, code, mathematics), and whether the grammatical-prior explanation is the right one.
Key Contributions
- The first systematic evaluation of pre-pretraining at both model and data scale, spanning 500M to 7B parameters, pre-training budgets up to 100B tokens, four pre-training data mixtures, and five PPT tasks.
- Evidence that the downstream performance and token efficiency benefits of PPT persist under parameter scaling and extended pre-training up to 100B tokens, saving at least 21B pre-training tokens at the 3B scale and remaining positive even at 7B.
- A negative result for the prevailing explanation: there is no consistent evidence that PPT acts as a grammatical prior, since downstream performance does not consistently align with grammatical acceptability.
- Evidence that PPT gains depend on the presence of web text rather than on the proportion of code and mathematics: a 13x increase in mathematical data leaves the downstream gain unchanged, whereas removing web text collapses it.
Main Findings
- Downstream gains persist across scales and mixtures. Of the 12 scale-mixture pairs (4 PT data mixtures x 3 model scales), 9 gain at least 0.6 points on the downstream average, with a mean gain of 1.6 among them. The gain is similar across scales: 0.8 at 500M, 1.4 at 1B, and 1.3 at 3B.
- A few pairs do not benefit. The remaining three pairs gain little or lose: C4 at 500M (+0.3), OLMo3 at 1B (−0.5), and OLMo3 at 3B (+0.0), leaving OLMo3 as the only mixture that fails to benefit at more than one scale.
- Gains come from synthetic data, not extra optimization. Control, which runs the same 500 warm-up steps on held-out PT text, stays close to PT-Only at every scale (mean differences of −0.2 at 500M, +0.4 at 1B, +0.2 at 3B). k-Shuffle Dyck outperforms Control in 10 of 12 pairs, with mean margins of 1.0, 1.0, and 1.1 points at the three scales.
- Gains are broad but strongest on context-dependent tasks. Each category improves in at least 9 of 12 pairs (reading comprehension 11, science QA 10, commonsense reasoning 9, language modeling 10). Language modeling gains the most at all three scales (1.2, 1.8, and 1.7 points, against 0.3 to 1.3 for other categories). ReCoRD and HellaSwag improve in all 12 pairs and LAMBADA in 10.
- No consistent grammatical-prior effect. k-Shuffle Dyck improves BLiMP accuracy in 9 of 12 pairs, but only one of these gains is stable (OLMo3 at 1B), against 10 of 12 for the downstream average. At 500M, results are mixed: OLMo3 (+0.6) and Marin (+0.3) improve, while SmolLM3 (−0.3) and C4 (−0.6) degrade — the largest degradation is on C4, where prior work reports gains. On Marin the delta moves from +0.3 at 500M to +1.6 at 1B and back to −0.3 at 3B. Neither the Morphology nor the Syntax subgroup shows a stable gain.
- Long-range retrieval improves instead. k-Shuffle Dyck lowers verbatim-retrieval NLL in all 12 scale-mixture pairs, with a stable reduction in 7 of them, against 1 of 12 for BLiMP. Control improves in 8 pairs but only 3 of those reductions are stable.
- PPT task type matters, but the formal/non-formal divide does not. At 3B, MP-Struct Core (+0.9 average) and NCA (+1.2) yield gains comparable to k-Shuffle Dyck (+1.3), and all three improve 3 of the 4 mixtures. Set degrades downstream performance by 7.2 points and improves none of the mixtures. The three effective tasks all require retrieving a specific earlier position; Set only requires tracking which tokens have been seen.
- Code and mathematics do not make PPT redundant. On Marin at 3B, raising math from 1.3% to 17.0% (lowering web share from 92.6% to 76.9%, matching OLMo3) gives an average gain of 2.0 points against 1.9 for the full Marin mixture. DCLM-only retains 1.5 of the 1.9 points and FineWeb-Edu-only retains 1.0. Removing web text entirely (leaving 82% source code and 18% mathematics) reduces the average PPT gain to 0.2 points.
- Gains survive long budgets. The downstream average improves in all 14 budget points across the three extended runs, by 1.0 on average on C4, 1.3 on Marin, and 0.6 at 7B, ending at +1.6 on C4 and +1.2 on Marin at 100B tokens, against +1.2 and +1.1 at 21B. On Marin, PPT reaches 62.3 on average after 63B tokens while PT-Only reaches 62.0 after 84B, saving at least 21B pre-training tokens.
- The 7B gain narrows but stays positive. On Marin the average gain falls from 1.3 at 3B to 0.6 at 7B, yet remains positive at every budget. The largest 7B gain occurs at the longest budget (+1.0 at 75.5B tokens).
- BLiMP stays mixed at scale. Across the 14 extended-run budget points, BLiMP improves in only 4 while verbatim retrieval improves in 10.
Methodology in Plain English
The researchers take a standard language-model pre-training run and insert a short warm-up phase before it. In that warm-up, a randomly initialized model trains for 500 steps on synthetic sequences — bracket-matching languages, marker-annotated structures, deduplication-style sequences, or neural cellular automaton states — and the resulting weights then initialize the real pre-training run, with optimizer moments and schedule reset. Because only the initialization state differs, any difference from a baseline trained from scratch can be attributed to the warm-up.
They vary four things. First, the warm-up task: two formal languages (k-Shuffle Dyck and MP-Struct Core), two structured synthetic tasks without formal grammars (Set and NCA), and an in-domain text control sampled from the pre-training corpus itself. Second, the pre-training data mixture: C4 (100% web), SmolLM3 (85.0% natural language, 79.8% web, 12.0% code, 3.0% math), OLMo3 (89.5% natural language, 76.9% web, 7.1% code, 3.4% math, plus 12.6% OCR-extracted academic PDFs), and Marin (92.6% natural language and web, 6.1% code, 1.3% math). Third, model scale: 500M, 1B, 3B, and 7B, all using the SmolLM3 architecture and tokenizer with a 128,256-token vocabulary and a 4,096-token context window. Fourth, pre-training budget: a standard 21B-token run (10K steps, batch size 512 sequences, 2.1M tokens per step) and an extended 100B-token run (47,684 steps) at 3B, plus a 7B run on Marin up to 75.5B tokens.
Evaluation uses ten downstream benchmarks in four groups (reading comprehension: RACE, ReCoRD; science QA: SciQ, ARC-Easy, OpenBookQA; commonsense reasoning: HellaSwag, PIQA, COPA, SocialIQA; language modeling: LAMBADA), plus two linguistic-competence tests: BLiMP for grammatical acceptability and a verbatim-retrieval task measuring mean NLL on repeating a noun list seen earlier in context. Models are evaluated every 1K steps, with mean and standard deviation reported over the second half of training (5K to 10K steps); a gain is called stable when its mean exceeds its SD across those checkpoints. Each configuration is a single run except PT-Only and k-Shuffle Dyck at 3B on Marin, which use three random seeds.
Why This Matters
The paper reframes a widely cited explanation for why synthetic pre-pretraining works. If the benefit is not a grammatical prior but a long-range retrieval capability, then the design of future warm-up tasks should target positional retrieval rather than mimic natural-language hierarchy. It also establishes that the technique is practically viable at scale: gains hold at 7B parameters and up to 100B pre-training tokens, and at 3B on Marin the approach saves at least 21B pre-training tokens for about 1% of the pre-training compute.
Real-world applications implied by the findings:
- Cheaper foundation-model training pipelines, where a short synthetic warm-up substitutes for tens of billions of pre-training tokens.
- More targeted data-mixture planning, since the results show PPT gains depend on including web text rather than on the code and math share of the mixture.
- Task-design guidance for anyone building synthetic curriculum data, favoring tasks that require retrieving a specific earlier position over tasks that only require tracking seen tokens.
- A more reliable evaluation template for claims about inductive biases, separating grammatical acceptability from context-dependent retrieval.
Industry relevance: pre-training is described in the paper as prohibitively expensive, often consuming tens of trillions of tokens. A warm-up phase costing about 1% of pre-training compute that saves at least 21B tokens at the 3B scale is directly relevant to organizations training models at these scales, and the negative result about grammatical priors is relevant to teams designing synthetic data curricula.
Future Directions
- Explaining OLMo3. The paper leaves a direct ablation of OLMo3, and the property responsible for its weak PPT gains, for future work. Neither the math share nor the web share explains the gap: the 17% Math variant matches OLMo3's 76.9% web share yet retains an average gain of 2.0 points, and SmolLM3 differs from OLMo3 by only three points in web share yet gains 1.8 points at 3B against 0.0 for OLMo3.
- Designing PPT tasks explicitly around long-range retrieval, since the evidence suggests that is the active ingredient rather than hierarchical structure.
- Determining whether the narrowing of gains at 7B reflects capacity or budget. The 7B runs reach roughly 11 tokens per parameter against 33 at 3B, and the largest 7B gain occurs at the longest budget.
- Investigating the web-text dependency, given that removing web text entirely reduces the average PPT gain to 0.2 points while the web-free condition is described as one no practical mixture approaches.
Target Audience
Researchers and engineers working on language-model pre-training, data curation, and training efficiency; practitioners designing synthetic data or curriculum warm-up phases; and anyone evaluating claims about inductive biases transferred from synthetic to natural language, particularly those who need the scale caveats (previous PPT work sat at or below 1B parameters and under 2B pre-training tokens) spelled out.
Authors’ abstract
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.