Research
MiniLingua: A Small Open-Source LLM for European Languages
Overview Research area: Natural Language Processing — multilingual large language model pre-training, tokenizer design, and instruction tuning. Technical level: Intermediate. The paper assumes familia

- arXiv
- 2512.13298
- Published
- 2025-12-15
- Authors
- Anna Aksenova, Boris Zverkov, Nicola Dainese, Alexander Nikitin, Pekka Marttinen
AI summary
Overview
Research area: Natural Language Processing — multilingual large language model pre-training, tokenizer design, and instruction tuning.
Technical level: Intermediate. The paper assumes familiarity with transformer architectures, tokenization, and standard LLM training stages (pre-training, SFT), but explains its design choices in prose accessible to readers who know the basics.
Scope: The paper describes the end-to-end development of MiniLingua, a one-billion-parameter open-source multilingual LLM trained from scratch for 13 European languages, covering data curation, tokenizer training, pre-training, supervised fine-tuning, and multilingual benchmark evaluation against EuroLLM-1.7B, Gemma-3-1B, and SmolLM2-1.7B.
What This Paper Is About
Most capable large language models are proprietary, expensive to run, and trained predominantly on English, which leaves many European languages underserved. MiniLingua addresses this by building a compact, fully open, roughly one-billion-parameter model trained from scratch for 13 European languages, with both base and instruction-tuned versions. The goal is to show that careful data curation, a balanced multilingual tokenizer, and efficient training under academic compute limits can produce a model competitive with larger-budget academic models.
Key Contributions
- Base and instruction-tuned open-source multilingual LLMs with approximately one billion parameters, trained from scratch for 13 European languages (English, German, Spanish, French, Italian, Dutch, Portuguese, Polish, Swedish, Czech, Finnish, Bulgarian, and Greek), released at
https://huggingface.co/minilingua-ai. - Fully open source code for data cleaning, pre-training, fine-tuning, and evaluation, with model weights and tokenizer released at
https://huggingface.co/minilingua-ai/MiniLingua-1b-Instruct. - Curated multilingual datasets with permissive licenses for research use, including a purpose-built multilingual question-answering supervised fine-tuning dataset (
https://huggingface.co/datasets/minilingua-ai/mcqa-minilingua-sft), plus released training artifacts and data (https://github.com/MiniLingua-ai/training_artifacts/tree/main/data). - A from-scratch multilingual tokenizer study covering four vocabulary sizes and four data mixtures, evaluated with Normalised Sequence Length, showing that mixture design can matter more than vocabulary size alone.
Main Findings
- Tokenizer beats larger-vocabulary competitors: The best MiniLingua tokenizer (128K vocabulary, Balanced mixture) achieved lower Normalised Sequence Length than both GPT-4o and EuroLLM across all data mixtures tested, and achieved the best compression for nearly all evaluated languages. Compression was notably better for Greek, Bulgarian, Finnish, and Czech, while remaining competitive for English (lower NSL than EuroLLM). Only code and English were better compressed by GPT-4o's much larger vocabulary.
- Data mixture mattered more than vocabulary size: The Balanced configuration (English at only 20% of tokenizer data, no language below 3%) showed strong performance even at 64K tokens, indicating that data mixture design can have a greater effect than vocabulary size alone.
- Base model is competitive despite far less compute: MiniLingua-1b scored COMET 0.343 on FLORES, 0.230 on Belebele, 0.248 on SIB, and 0.240 on MMLU, using 1.5T training tokens and 75K GPU hours. It outperformed Gemma-3-1b-pt and EuroLLM-1.7b on SIB-200 and MMLU-X, and approached EuroLLM on FLORES (0.360). SmolLM2-1.7b remained strongest overall, attributed to its 11T-token training set.
- Instruction-tuned model leads on open-ended generation: MiniLingua-1b-Instruct reached 0.681 on FLORES COMET and 0.187 ROUGE on MassiveSumm, the highest in the instruction-tuned comparison. Gemma-3-1b-it and SmolLM2-1.7b-Instruct scored 0.494 / 0.001 and 0.496 / 0.015 respectively, with the paper attributing their low summarization scores to failing to consistently output the target language and defaulting to English.
- Instruction-tuned model beats EuroLLM across all tasks: Against the similarly trained but larger-budget EuroLLM-1.7b-Instruct, MiniLingua scored higher on Belebele (0.262 vs 0.216), SIB (0.149 vs 0.124), MMLU (0.245 vs 0.205), and MSum (0.187 vs 0.0138); EuroLLM's FLORES result is omitted because it was trained on that dataset's dev set.
- Weakness on strict format tasks: On Belebele, SIB, and MMLU, MiniLingua underperformed RLHF-aligned models such as Gemma-3-1b-it (0.366 / 0.558 / 0.311) and SmolLM2-1.7b-Instruct (0.369 / 0.613 / 0.271), largely due to incorrect output format generation.
- Decay-stage data mixture effect: Increasing the high-quality data share to 30% while keeping the pre-training language distribution gave the best validation loss (2.16), compared to 2.30 at 10% with the pre-training distribution and 2.17 at 30% with upsampled low-resource languages.
- English remains the largest single share but stays below 40%: In the final data distribution, English never exceeds 40% of the data, compared to 50% for EuroLLM.
Methodology in Plain English
Data. The team assembled three data groups: general web text from the FineWeb-2 corpus (derived from 96 Common Crawl snapshots spanning summer 2013 through April 2024, about 8 terabytes of compressed text and nearly 3 trillion words); manually collected high-quality sources such as news, books, subtitles, encyclopedias, government documents, and speech transcripts; and code from The Stack v2. A six-step cleaning pipeline applied language filtering, heuristic filters, repetition filters, an inappropriate-term blacklist, near-duplicate removal, and cross-deduplication against evaluation sets using Jaccard similarity. FineWeb-2 received only the obscenity filter and evaluation deduplication because it was already rigorously cleaned; the high-quality data went through the whole pipeline; the code data went through only the final step. The team aimed to retain 80% of collected data per language.
Tokenizer. Four vocabulary sizes (32K, 64K, 96K, 128K) were trained on cleaned FineWeb subsets under four data distributions (Balanced, Intermediate, Original, Train). Sizes were kept divisible by 8 for hardware compatibility. Evaluation used Normalised Sequence Length — the token-to-word ratio, where lower is better — reported as average NSL (per-language average) and weighted NSL (weighted by dataset size).
Small-scale experiments and scaling law. Before the main run, the team trained three smaller models of roughly 30M, 60M, and 110M non-embedding parameters, architecturally aligned with the target model, to fit a bi-variate power-law describing loss as a function of model size and dataset size. The fit used nonlinear least squares (scipy.optimize.curve_fit) with constrained parameter ranges and up to 30,000 function evaluations.
Hyperparameter search. Learning rate and batch size were tuned on small models, exploring global batch sizes from 32K to 2M tokens (1 to 32 nodes) and learning rates of 0.0005, 0.001, 0.0025, 0.004, and 0.005, measured by validation loss after a fixed budget. A global batch size of 2M tokens and a learning rate of 5×10⁻⁵ were reported as selected for the stable stage; the later training-run description states the stable stage learning rate was held constant at 0.0025.
Pre-training. MiniLingua is a decoder-only transformer with 32 layers, hidden size 1536, SwiGLU feed-forward size 6144, 24 attention heads with 6 key-value heads (grouped query attention), rotary positional embeddings, RMSNorm, tied input-output embeddings, and a 128,008-token vocabulary. Training ran about 12 days on 32 nodes with four AMD MI200X GPUs each on the LUMI supercomputer, using Megatron-LM optimized for AMD architecture, over 715,000 iterations. The Warmup-Stable-Decay schedule linearly increased the learning rate to 0.0025 over 6,000 steps, held it for 40,000 steps, reduced it to 0.0005, then decayed it to zero over the final 10% of training. That decay stage used 30% high-quality data with the remaining share distributed as in the stable stage.
Supervised fine-tuning. Training combined public instruction datasets (Aya, Bactrian-X, LIMA, and Self-OSS-Instruct Code 50k) with a curated multilingual QA dataset generated with GPT-4o, varying instruction language (20% English, 80% target language), instruction wording, expected answer format, and output formatting. Instances were wrapped in special tokens and filtered for length, triviality, and language-tag mismatches. Three sampling strategies were compared (Original, Train, and Equal distribution); the Equal distribution was chosen. The final run lasted about 25 hours on four H200 GPUs, one epoch, maximum effective batch size 256, learning rate 2×10⁻⁵, with loss computed only on response tokens.
Evaluation. Base and instruction-tuned models were compared on FLORES-200 (3,001 parallel sentences), Belebele (122 languages), MMLU-X (57 knowledge domains), and SIB-200, with MassiveSumm (28.8M news articles, 92 languages) added for instruction-tuned models. Base models were scored via log-probabilities of candidate answer tokens; instruction-tuned models were scored by fuzzy string matching of generated outputs against expected answers. Open-ended tasks used COMET for translation and ROUGE for summarization.
Why This Matters
Impact on research. The paper demonstrates that multilingual balance and tokenizer data-mixture design — rather than raw scale — can drive competitive results within academic compute limits, and it releases models, tokenizer, code, training curves, configuration files, evaluation scripts, and permissively licensed datasets. That combination of a documented one-billion-parameter multilingual model and open artifacts supports reproducibility work that is difficult with proprietary or partially closed models.
Real-world applications:
- On-device assistants and translation tools for European languages, since the model is small enough for constrained hardware.
- Summarization, classification, and question answering in lower-resource European languages such as Greek, Bulgarian, Finnish, and Czech.
- Privacy-sensitive deployments where sending text to a remote proprietary API is undesirable.
- Research and teaching environments where permissive licensing and available training code matter.
Industry relevance. The paper positions itself against models built with far larger budgets — Gemma-3-1B (trained on an undisclosed 2T-token dataset with distillation and RLHF) and SmolLM2-1.7B (11T tokens with multi-stage pre-training and RLHF) — showing where an academic, non-RLHF-aligned model wins (open-ended generation in the target language) and where it loses (strict single-letter answer formats). The released multilingual QA dataset, built by translating and rephrasing existing QA sources with GPT-4o, is directly reusable for fine-tuning other multilingual models.
Future Directions
- Systematic ablations: The authors state they did not run systematic ablations on data cleaning strategies or tokenizer vocabulary size because both were fixed early under compute constraints; these remain open questions.
- An alignment stage: The paper notes the absence of RLHF due to a lack of high-quality multilingual preference data, leaving alignment methods for small multilingual models unexplored.
- Better evaluation: The authors flag that their evaluations relied largely on translated, English-centric benchmarks, which may introduce cultural bias; native-language and culturally grounded benchmarks are needed.
- Broader language coverage and fit quality: Only 13 European languages are covered out of more used in Europe, and the scaling-law fit was limited by a narrow model-size range and few data points, which the authors say could be improved with larger models and longer training schedules.
Target Audience
This paper is most useful to multilingual NLP researchers and engineers building or fine-tuning small language models, particularly those working under constrained compute budgets; to practitioners who need an openly licensed, on-device model covering European languages; and to readers interested in tokenizer design, data-mixture selection, and reproducible training pipelines. It is also relevant to teams seeking permissively licensed multilingual instruction-tuning data.
Authors’ abstract
Large language models are powerful but often limited by high computational cost, privacy concerns, and English-centric training. Recent progress demonstrates that small, efficient models with around one billion parameters can deliver strong results and enable on-device use. This paper introduces MiniLingua, a multilingual open-source LLM of one billion parameters trained from scratch for 13 European languages, designed to balance coverage and instruction-following capabilities. Based on evaluation results, the instruction-tuned version of MiniLingua outperforms EuroLLM, a model with a similar training approach but a larger training budget, on summarization, classification and both open- and closed-book question answering. Moreover, it remains competitive with more advanced state-of-the-art models on open-ended generation tasks. We release model weights, tokenizer and source code used for data processing and model training.