Research
AfriNLLB: Efficient Translation Models for African Languages
Overview Research area: Natural Language Processing / Machine Translation (low-resource African languages, model compression, knowledge distillation). Technical level: Intermediate. The paper assumes

- arXiv
- 2602.09373
- Published
- 2026-02-10
- Authors
- Yasmin Moslem, Aman Kassahun Wassie, Amanuel Gizachew Abebe
AI summary
Overview
Research area: Natural Language Processing / Machine Translation (low-resource African languages, model compression, knowledge distillation).
Technical level: Intermediate. The paper assumes familiarity with encoder–decoder NLLB-style translation models, standard MT metrics (BLEU, chrF++, COMET), and compression techniques (layer pruning, float16 quantization), though each is explained enough for a motivated reader.
Scope: The paper introduces AfriNLLB, a family of compressed NLLB-200 600M variants covering 15 language pairs (30 translation directions) across 10 African languages plus 5 African Union official languages, released alongside curated training data, and shows that iterative decoder-layer pruning plus fine-tuning preserves translation quality while substantially increasing inference throughput.
What This Paper Is About
Translation resources for African languages are scattered across many sources and are expensive to collect, and most African languages remain low-resource despite having millions of speakers. AfriNLLB addresses this by curating parallel corpora for African languages and then compressing the NLLB-200 600M translation model using iterative layer pruning, quantization, and knowledge distillation, so that strong translation quality can run on resource-constrained hardware.
Key Contributions
- A curated, filtered training corpus for African languages. The authors collect data from OPUS, Hugging Face, GitHub, and other public sources, and run it through a four-stage pipeline (rule-based filtering, language detection, semantic filtering, quality estimation), taking 7,184,998 raw pairs down to 6,443,951 processed pairs and 1,609,411 sampled pairs for 15 language pairs. All training data used to fine-tune the baseline and pruned models is released.
- A family of compressed AfriNLLB models. Built on NLLB-200 600M, the models are produced via iterative decoder-layer pruning (removing the least important layer one at a time, guided by chrF++ on the Flores200 dev split), followed by re-fine-tuning and sequence-level knowledge distillation from an NLLB-200 3.3B teacher.
- Multiple released configurations. Models with 12 encoder / 8 decoder layers (548M), 12/6, 12/4, and 8/8 (498M) layer configurations are released in two forms: a Transformers version that supports further fine-tuning and a CTranslate2 version for efficient inference.
- An ablation study of pruning strategies. The paper compares iterative versus middle-layer pruning (layers 4–7), decoder-only versus encoder+decoder pruning, and 4, 6, and 8 removed decoder layers.
Main Findings
- Compression preserves quality on average. Across all 30 directions, the AfriNLLB 548M model (iterative pruning + fine-tuning) reaches an average BLEU of 27.05, chrF++ of 51.41, and COMET of 63.65, versus 26.21 / 49.93 / 63.31 for the NLLB-200 600M baseline — gains of 3.2%, 3.0%, and 0.54% respectively, as reported in Table 7.
- Throughput improves substantially. The same table reports average output throughput of 1833.95 tokens/second for the 548M model versus 1468.2 for the baseline (+24.9%), rising to 3532.15 tokens/second with float16 quantization (+140.57%).
- Speed gains reported at several granularities. Section 4 states that pruning 4 decoder layers created a 548M model that is "23% faster in average than the baseline." The Table 4 caption describes pruned models as "up to 20% faster than the baseline without quantization, and 57% faster with float16 quantization." The conclusion refers to "over 20–50% inference performance gains."
- Direction-level averages hold up. In Table 4, for xx→en the pruned+FT model scores 34.01 BLEU / 56.98 chrF++ / 71.20 COMET versus 35.15 / 57.61 / 71.87 for the fine-tuned baseline; for en→xx it scores 24.17 / 50.05 / 70.37 versus 24.28 / 49.97 / 70.91.
- The largest per-pair gains are in low-resource directions. Table 7 shows English→Yoruba BLEU rising from 4.32 to 8.03 (+85.9%) and chrF++ from 22.87 to 29.20 (+27.7%); English→Wolof BLEU from 5.06 to 6.97 (+37.7%); English→Egyptian Arabic from 11.87 to 14.94 (+25.9%); English→Swahili and English→Hausa each +16.7%.
- Some directions regress. Table 7 shows English→Somali BLEU falling from 11.38 to 11.05 (−2.9%), English→Spanish from 26.71 to 24.78 (−7.2%), English→Portuguese from 46.45 to 42.72 (−8.0%), and Egyptian Arabic→English from 30.64 to 28.69 (−6.4%).
- COMET degrades sharply for the French–Lingala/Wolof group. In Table 4, xx→fr COMET drops from 18.47 (fine-tuned baseline) to 14.52 (pruned+FT) and falls further in Table 5 to 11.78 for the 12/6 configuration and 5.67 for the 12/4 configuration.
- Iterative pruning beats middle pruning. The ablation finds that iterative layer pruning clearly outperforms middle layer pruning for both the decoder-only and encoder+decoder cases, and that fine-tuning after pruning is "crucial in all cases."
- Quality retention is tolerant up to a point. The ablation states that removing up to 6 decoder layers (50%) yields similar or better performance compared to the NLLB-200 600M baseline, thanks to fine-tuning before and after pruning.
- Encoder-layer removal remains ambiguous. The paper states it "is not clear to what extent" removing encoder layers affects quality, and notes that keeping encoder layers intact had been recommended in prior speech work.
- Data processing choices. A filter threshold of 0.6 was chosen after experimenting with different values; semantic filtering was skipped for Lingala because no supporting embedding model was found.
Methodology in Plain English
The team started from NLLB-200 600M, a multilingual translation model, and fine-tuned it on parallel data they assembled for African languages. To build that data, they pulled parallel corpora from OPUS, Hugging Face, GitHub, and other public sources, then cleaned it in four stages: rule-based cleanup (deduplication, dropping empty segments, removing HTML tags, excluding very short or very long sentences and pairs with a source-to-target length ratio above 2x), language identification (AfroLID for African languages, fastText for others), semantic filtering with sentence embeddings (LaBSE for African languages, DistilUSE for high-resource pairs) to check that source and target mean similar things, and reference-free quality estimation (AfriCOMET-QE-STL for African languages, COMET for high-resource pairs). After merging, deduplication, and downsampling high-resource pairs to 200,000 each, the training set totals 1.6M samples, or 3.2M bidirectional samples after reversing.
For compression, they used a greedy procedure: test how much translation quality drops when each decoder layer is removed, prune the single least harmful layer, and repeat until reaching the target (4, 6, or 8 layers). chrF++ measured on the Flores200 dev split, mostly with African languages as the target, drove these decisions. Because pruning hurts quality, each pruned model was fine-tuned for one epoch (learning rate 5e-5, batch size 8, gradient accumulation 4, early stopping with patience 10, evaluated every 1000 steps, on a single A40 48GB GPU). They also used sequence-level knowledge distillation, training the models on a mix of authentic data and 568k segments of synthetic data generated by a larger NLLB-200 3.3B teacher. Evaluation used BLEU and chrF++ from sacreBLEU, AfriCOMET (africomet-mtl) for African languages, and COMET (wmt22-comet-da) for Arabic and European languages, with validation on the Flores200 dev split (997 segments) and testing on the devtest split (1,012 segments); inference used CTranslate2 with beam size 3 and a batch size of 1024 tokens on an A40 48GB GPU.
Why This Matters
The work targets a practical bottleneck: African languages have millions of speakers but few datasets and models, and deploying large multilingual models is costly. By showing that a 600M-parameter model can be compressed to roughly 548M (or 498M) while keeping comparable quality and running markedly faster, the paper offers a template for building deployable translation systems under hardware constraints.
Real-world applications:
- Public services: translation support in government and health settings, which the paper explicitly identifies as a challenge for speakers of low-resource African languages.
- Health and legal document translation: several of the source datasets in Table 6 are health-related (for example, AfriDocMT-health and ELRC-wiki_health), suggesting direct applicability to medical content.
- On-device or edge deployment: the CTranslate2 + float16 releases suit environments without large GPUs.
- Further model development: the released Transformers-format models and training data let others fine-tune for additional languages.
Industry relevance: The released compression recipe (iterative pruning guided by a quality metric, followed by fine-tuning and distillation) is directly transferable to other encoder–decoder systems where inference cost matters. The dual release format — a trainable Transformers version and an optimized CTranslate2 version — maps onto the split between research iteration and production serving.
Future Directions
- Expand language coverage: the authors state that future versions of AfriNLLB will add more languages.
- Additional data augmentation: beyond knowledge distillation, they plan to investigate techniques such as back-translation.
- Other architectures: they intend to expand the approach to autoregressive large language models and encoder-only models.
- Encoder layer pruning: the paper explicitly leaves open how much removing encoder layers affects quality and whether the recommendation to keep encoders intact, drawn from prior speech work, applies to text encoder–decoder models like NLLB-200. They also note a plan to experiment with using both translation directions during layer importance evaluation.
Target Audience
Machine translation researchers and engineers working on low-resource languages; the African NLP research community; practitioners who need to deploy translation models on limited hardware; and anyone interested in model compression methods (layer pruning, quantization, knowledge distillation) applied to large multilingual models. Readers seeking a beginner-level introduction to MT will find the data pipeline and evaluation setup accessible, while the pruning and distillation details require some prior background.
Authors’ abstract
In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 language pairs (30 translation directions), including Swahili, Hausa, Yoruba, Amharic, Somali, Zulu, Lingala, Afrikaans, Wolof, and Egyptian Arabic, as well as other African Union official languages such as Arabic (MSA), French, Portuguese, and Spanish. Our training data covers bidirectional translation between English and 13 languages, and between French and two languages (Lingala and Wolof). AfriNLLB models are based on NLLB-200 600M, which we compress using iterative layer pruning and quantization. We fine-tune the pruned models on parallel corpora we curated for African languages, employing knowledge distillation from a larger teacher model. Our work aims at enabling efficient deployment of translation models for African languages in resource-constrained settings. Our evaluation results demonstrate that AfriNLLB models achieve performance comparable to the baseline while being significantly faster. We release two versions of the AfriNLLB models, a Transformers version that allows further fine-tuning and a CTranslate2 version for efficient inference. Moreover, we release all the training data that we used for fine-tuning the baseline and pruned models to facilitate further research.