Research
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models Overview Research area: Natural Language Processing / AI for mathematics — specifically, extracting latent mathematical knowledge

- arXiv
- 2609.32264
- Published
- 2026-09-26
- Authors
- Pavel Tikhonov, Elena Tutubalina, Ivan Oseledets, Dmitry I. Ignatov, Mikhail Seleznyov
AI summary
LANTERN: Illuminating Hidden Mathematical Knowledge in Language ModelsOverview
Research area: Natural Language Processing / AI for mathematics — specifically, extracting latent mathematical knowledge from the internal activations of a pretrained large language model to discover undocumented relations between mathematical objects.
Technical level: Intermediate. The paper is readable without specialist background, but it assumes familiarity with ideas such as hidden-state activations, linear classifiers, ROC-AUC, and generating functions.
Scope: The authors build LANTERN, a pipeline that ranks pairs of On-Line Encyclopedia of Integer Sequences (OEIS) entries using a classifier trained on model activations, then filters, generates, and verifies candidate relations, yielding 62 technically verified relations from 50 million candidate pairs in under 8 hours.
What This Paper Is About
Language models can now prove theorems and solve olympiad problems, but a human still decides which problems are worth attacking. The paper asks whether a model's internal representations already encode knowledge of mathematical relationships that no one has written down, and whether that knowledge can be mined directly. The authors target the OEIS — a versioned encyclopedia of integer sequences with explicit cross-references — because it offers well-defined objects, documented relations, and cheap verification of exact claims about sequences.
Key Contributions
-
Evidence that pretrained models encode undocumented mathematical relations. The authors show that a linear classifier over Qwen3-32B hidden states can rank which OEIS entries are related, using only the one-line definitions of the entries as input.
-
LANTERN, a cost-effective discovery pipeline. The method combines classifier ranking, staged filtering with cheap and stronger agents, executable verification against stored sequence terms, analytical checking, and a post-verification content screen.
-
62 technically verified relations between previously unlinked OEIS sequences. These came from 50 million pairs ranked across 10,000 frequently referenced entries. A content screen retained 13 relations, of which nine were graded informative or insightful.
-
Four relations not found in the OEIS or in the authors' targeted literature search. Two are informative (S1) and two are insightful (S2), including two cross-domain bridges between cellular automata and between analytic q-series and 5-core partitions.
Main Findings
-
The funnel result: 62 – 13 – 9. The main run narrowed 50 million pairs to 500 ranked candidates, of which 118 passed the inexpensive filter, 44 received a hypothesis from the stronger agent, and all 44 passed sandbox verification and analytical checking. Across three configurations the pipeline produced 62 verified statements — 44 from the main configuration and 18 from two variants — and the content screen retained 13, of which nine were informative or insightful.
-
Two novel cross-domain bridges. One connects the one-dimensional Rule 150 cellular automaton (A071053) to the two-dimensional Fredkin replicator (A160239) via the identity g(k) = f(k)² − f(k−2)²; the other connects the Rogers–Ramanujan continued fraction (A007325) with a theta quotient (A227216) through T(q)²R(q)⁵ = Σ c₅(n)qⁿ, which counts 5-core partitions. Both were graded N2/S2 — not found in the OEIS or in the targeted literature search, and structurally insightful.
-
The block identity underlying the automata bridge was already recorded. The individual counting laws were known, and the block identity itself is recorded in OEIS A246030 (Conrad, 2023) in terms of Jacobsthal numbers; what was not recorded is its reading as a bridge between the two automata.
-
Activations carry more than the definitions do. On 793 held-out cross-references with a matched topical background, a surface-text scorer built on character 3–5-gram TF-IDF features reduced to 128 LSA dimensions was at chance (AUC = 0.51), while the activation-based classifiers reached AUC = 0.57–0.58 across all three model sizes. On the 42 substantive pairs the values were 0.61 and 0.67–0.69.
-
The classifier does not need to have seen the pages. On 2,379 entries created in 2026 after the model's knowledge cutoff, their 5,557 cross-references to older entries were scored by a classifier trained only on links among older entries; it reached AUC = 0.88, compared with 0.89 on a random held-out split among older entries. The corresponding text scorer dropped from 0.85 to 0.78.
-
Asking the model directly works about as well as the trained classifier. A label-free readout — showing Qwen3-32B two definitions and asking it to rate the depth of their relation on a 0–9 scale — achieved a median percentile of 88.5 on the 74 substantive future pairs, against 88.0 for the classifier. The classifier's advantage is throughput, not accuracy.
-
The pipeline is fast. The entire end-to-end process including classifier training, candidate ranking, filtering and verification took under 8 hours.
-
Label composition. Of the 32,263 cross-referenced pairs examined, only gold and silver graded links — 2,780 pairs, or 9% — were used as positives; the rest were treated as negatives or excluded.
Methodology in Plain English
The authors picked the OEIS because every entry has a one-line definition, a list of initial terms, and a record of which other entries link to it. They took the 10,000 most-frequently-mentioned sequences from a 2026-08-15 snapshot (the corpus was built from S2; other snapshots were S0 of 2025-04-30 and S1 of 2026-02-17) and formed all 50 million pairs.
Labeling. Naively treating "is cross-referenced" as the label risks teaching the classifier to recognize trivial links. So they had Claude Sonnet 5 grade every linked pair as gold (a theorem-grade identity, bijection, asymptotic or density result, or stated conjecture), silver (real but routine or textbook content), or trivia (a rescaling, shift, or template twin). Every gold verdict was re-read in full by Claude Fable 5.1, and 91–94% held. Gold and silver — 2,780 pairs — became the positives, with trivia links, bare cross-references, and random unlinked pairs as negatives.
Embedding. Each definition was passed once through Qwen3-32B, and the residual stream was kept after every layer and averaged over tokens, giving 65 layer vectors of width 5,120 per entry. The classifier does not see raw embeddings, which would overwhelm 2,780 positives. Instead each pair is reduced to 321 features: a centred cosine at every layer (subtracting each entry's typical similarity to a sample of 512 corpus entries), plus the element-wise product and absolute difference of PCA-reduced vectors at two layers chosen at the same relative depth (layers 4 and 40 for the 65-layer models). A logistic regression maps these to a score, trained on a rebalanced set capped at eight positives per entry and fifty per contributor.
From scores to verified statements. Four stages: score all 50 million pairs and walk down the ranking keeping only pairs where neither entry already appeared (the top 500 for the main configuration); pass them through Claude Sonnet 5 without tools as a cheap plausibility filter; hand survivors to Claude Opus 5, which has a shell, a local OEIS mirror, and a library for exact arithmetic of truncated power series, and must return a derivation, a proof class, and a Python function; then verify — the function runs in a sandbox with no files, network, or third-party libraries and must match every stored term (25 to 102 per pair), after which the derivation is checked analytically by hand.
Content screening. Technical verification is necessary but not sufficient, since a relation can reproduce every stored term while being mathematically empty. A post-verification screen excluded notation and parameter changes, elementary index or value transformations, direct consequences of a definition, generic reconstruction recipes, finite-value coincidences, and redundant members of repeated families. Survivors were then labelled on two independent axes: significance (S0 routine, S1 informative, S2 insightful) and novelty (N0 documented in the OEIS, N1 known elsewhere, N2 not found in the OEIS or the targeted literature search).
Controls. Three alternative explanations were tested: that surface text suffices, that the model memorized OEIS pages, and that the signal comes from the authors' labels rather than the model.
Why This Matters
The paper reframes the role of language models in mathematics from answering questions to choosing them, which the authors call the bottleneck of AI-assisted mathematics. It also demonstrates that a knowledge base's own structure can be mined for missing edges using nothing but a model's hidden states and a logistic regression, at a cost of hours rather than months.
Real-world applications:
- Maintaining and enriching large structured knowledge bases. The same recipe — train on existing links, rank unlinked pairs, verify executably — applies to biological, chemical, or bibliographic databases where links are recorded but incomplete.
- Speeding up expert curation of the OEIS and similar encyclopedias. The OEIS has been built by mathematicians over decades, and extending a single entry can take decades; automated candidate generation with verification narrows what experts need to read.
- Guiding research direction. The two cross-domain bridges illustrate how a relation nobody recorded can suggest that two objects from different domains share a common specification.
- Literature-based discovery. The method is a computational successor to Swanson-style discovery, which predicted unrecorded connections between known concepts decades ago.
Industry relevance: The pipeline uses a single embedding pass over a corpus to produce scores for all pairs, which makes it a throughput-friendly component in any retrieval or knowledge-graph-completion system. The finding that activations outperform surface text on matched-background pairs is directly relevant to anyone building similarity or link-prediction services over technical documents.
Future Directions
- Does the approach transfer beyond the OEIS? The limitations section flags that the OEIS is unusually structured — definitions, formulas, initial terms, and explicit links — and that the corpus is biased toward frequently mentioned sequences. Whether the method works for less formal mathematical objects, or for knowledge bases that do not record relations as explicit links, is untested.
- Broader model coverage. The authors evaluated only one family of pretrained models (Qwen3-32B, Qwen3-8B, and Qwen3.5-27B). Whether the latent signal appears in other model families is unknown.
- Comparing against other candidate-selection strategies. The paper positions LANTERN against systems that take candidates from the literature, from knowledge-graph topology, or from the model's own generation, but does not report a head-to-head comparison on the same corpus.
- Establishing priority for the novel relations. The N2 label means "not found by the OEIS and targeted literature searches," and the authors explicitly state it does not establish priority — resolving whether the four novel connections are genuinely new theorems requires expert review that the paper does not report.
Target Audience
Researchers in AI for mathematics and automated scientific discovery, machine learning practitioners interested in probing latent knowledge in language models, and mathematicians or database curators interested in computational tools for finding unrecorded connections between known objects. Readers who want the technical mechanism will find the classifier design and the control experiments most useful; readers interested in outcomes can focus on the two worked discoveries and the 62–13–9 funnel.
Authors’ abstract
Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a classifier over pretrained-model activations to rank candidate relations, followed by staged filtering, hypothesis generation, executable verification, and analytical checking. Applied to the On-Line Encyclopedia of Integer Sequences (OEIS), LANTERN ranked 50 million pairs among 10,000 frequently referenced sequences and produced 62 verified relations between pairs without an existing OEIS cross-reference. A content screen retained 13 relations worth presenting; nine of these are informative or insightful, including four which are entirely novel to the best of our knowledge: none appears in the OEIS or in our targeted literature search. The entire end-to-end process including classifier training, candidate ranking, filtering and verification took under 8 hours.