Skip to content
AI.info

Research

CoFrGeNet: Continued Fraction Architectures for Language Generation

Overview Research area: Natural Language Processing, specifically neural architecture design for large language models — replacing components of the Transformer block with an alternative function clas

CoFrGeNet: Continued Fraction Architectures for Language Generation
arXiv
2601.21766
Published
2026-01-29
Authors
Amit Dhurandhar, Vijil Chenthamarakshan, Dennis Wei, Tejaswini Pedapati, Karthikeyan Natesan Ramamurthy, Rahul Nair

AI summary

Overview

  • Research area: Natural Language Processing, specifically neural architecture design for large language models — replacing components of the Transformer block with an alternative function class.
  • Technical level: Advanced. The paper assumes familiarity with Transformer blocks, causal attention, feed-forward networks, backpropagation, and continued-fraction mathematics (continuants, convergents, poles of rational functions).
  • Scope: The paper introduces CoFrGeNets (Continued Fraction Generative Networks), derives continued-fraction replacements for attention and feed-forward layers, derives closed-form gradients via continuants, and pre-trains GPT2-xl (1.5B) and Llama3 (3.2B) variants to compare against the originals on GLUE tasks, perplexity benchmarks, and zero-shot Q&A/reasoning tasks.

What This Paper Is About

Transformers are the dominant architecture for language generation, but their multi-head attention and feed-forward network (FFN) blocks are parameter-heavy. This paper asks whether a fundamentally different function class — continued fractions, previously used only for supervised learning in work by Puri et al. (2021) — can serve as a plug-in replacement for those blocks in a generative, causal setting. The goal is to match or beat standard Transformer performance on downstream tasks while using fewer parameters and less pre-training time.

Key Contributions

  1. Novel continued-fraction architectures for causal attention and FFNs. The authors propose two attention replacements — CAttnU (transposes the sequence-length dimension to force token mixing, with upper-triangular layers preserving causality) and CAttnM (no transpose; ladders map to attention weights computed causally over prior tokens) — plus Cffn, a gated, non-expanded (α = 1) FFN replacement. A model with both components replaced is called a CoFrGeNet; CoFrGeNet-F replaces only the FFN, CoFrGeNet-A only the attention.

  2. A continuant-based reformulation with custom gradients. The continued fraction is expressed as a ratio of continuant polynomials, and the paper proves a closed-form gradient (Proposition 1): ∂f̃(a)/∂a_k = (−1)^k (K_{d−k}(a_{k+1},…,a_d) / K_d(a_1,…,a_d))². This reduces the number of divisions from d (one per layer of a depth-d ladder) to a constant of 1, implemented as a custom torch.autograd.Function.

  3. A custom dyadic training schedule. Only the linear component is updated at first; higher-depth parameters are frozen and progressively unfrozen — depth i parameters are updated for t/2^i iterations out of a total of t iterations.

  4. Large-scale pre-training experiments. GPT2-xl (1.5B) pre-trained on OpenWebText and GneissWeb 35B, and Llama3 (3.2B) pre-trained on the docling data mix of 2T tokens spanning nine datasets (DCLM2, DCLM3Plus, FineWeb-2-edu, Starcoder, stack-edu, Finemath, Infiwebmath, opc-fineweb-math-corpus, Cosmopedia).

Main Findings

  • Parameter efficiency: CoFrGeNet-F uses 985M parameters, CoFrGeNet-A 1.21B, and the full CoFrGeNet 798M, versus GPT2-xl's 1.5B — i.e., roughly two-thirds to one-half the parameters. For Llama3, the variants are 2.1B (CoFrGeNet-F), 2.5B (CoFrGeNet-A), and 1.8B (full) versus Llama's 3.2B.

  • GLUE fine-tuning on OpenWebText: CoFrGeNet-F (985M) was best on all eight GLUE tasks, beating GPT2-xl on MNLI (87.26 vs 86.89), QQP (89.95 vs 88.93), QNLI (91.89 vs 91.35), SST2 (94.16 vs 93.56), COLA (82.59 vs 81.78), MRPC (80.21 vs 79.83), RTE (61.35 vs 60.27), and WNLI (58.30 vs 58.28).

  • GLUE fine-tuning on GneissWeb: CoFrGeNet-F again led on most tasks — MNLI 79.62 vs 78.28, QQP 87.26 vs 86.83, SST2 92.36 vs 91.82, COLA 74.83 vs 74.18, MRPC 78.01 vs 77.72, RTE 61.35 vs 60.19 — while GPT2-xl was best on QNLI (82.93 vs CoFrGeNet-F's 82.73). The GneissWeb-trained GPT2-xl at 1.5B also had the top COLA result among some variants at 74.18.

  • Perplexity on OpenWebText: CoFrGeNet-F was best on all six datasets — PTB 29.89 vs 30.12, Wikitext2 17.12 vs 18.30, Lambada 8.12 vs 8.66, AgNews 35.72 vs 37.13, LM1B 40.14 vs 41.20, Wikitext103 16.14 vs 17.50 (GPT2-xl values).

  • Perplexity on GneissWeb: CoFrGeNet-A was best on PTB (28.89 vs GPT2-xl's 29.07), while CoFrGeNet-F was best on Wikitext2 (18.13 vs 19.12), Lambada (30.52 vs 31.78), AgNews (41.63 vs 45.62), LM1B (46.83 vs 52.36), and Wikitext103 (18.11 vs 18.93).

  • **Continuant

Authors’ abstract

Transformers are arguably the preferred architecture for language generation. In this paper, inspired by continued fractions, we introduce a new function class for generative modeling. The architecture family implementing this function class is named CoFrGeNets - Continued Fraction Generative Networks. We design novel architectural components based on this function class that can replace Multi-head Attention and Feed-Forward Networks in Transformer blocks while requiring much fewer parameters. We derive custom gradient formulations to optimize the proposed components more accurately and efficiently than using standard PyTorch-based gradients. Our components are a plug-in replacement requiring little change in training or inference procedures that have already been put in place for Transformer-based models thus making our approach easy to incorporate in large industrial workflows. We experiment on two very different transformer architectures GPT2-xl (1.5B) and Llama3 (3.2B), where the former we pre-train on OpenWebText and GneissWeb, while the latter we pre-train on the docling data mix which consists of nine different datasets. Results show that the performance on downstream classification, Q\& A, reasoning and text understanding tasks of our models is competitive and sometimes even superior to the original models with $\frac{2}{3}$ to $\frac{1}{2}$ the parameters and shorter pre-training time. We believe that future implementations customized to hardware will further bring out the true potential of our architectures.

Read the original paper