Skip to content
AI.info

Research

The PLLuM Instruction Corpus

Overview Research area: Natural Language Processing — instruction dataset design for fine-tuning large language models, with a focus on Polish language and cultural adaptation. Technical level: Interm

The PLLuM Instruction Corpus
arXiv
2511.17161
Published
2025-11-21
Authors
Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński, Anna Cichosz, Zuzanna Deckert, Monika Garnys, Izabela Grabarczyk, Wojciech Janowski, Sylwia Karasińska, Aleksandra Kujawiak, Piotr Misztela, Maria Szymańska, Karolina Walkusz, Igor Siek, Maciej Chrabąszcz, Anna Kołos, Agnieszka Karlińska, Karolina Seweryn, Aleksandra Krasnodębska, Paula Betscher, Zofia Cieślińska, Katarzyna Kowol, Artur Wilczek, Maciej Trzciński, Katarzyna Dziewulska, Roman Roszko, Tomasz Bernaś, Jurgita Vaičenonienė, Danuta Roszko, Paweł Levchuk, Paweł Kowalski, Irena Prawdzic-Jankowska, Marek Kozłowski, Sławomir Dadas, Rafał Poświata, Alina Wróblewska, Katarzyna Krasnowska-Kieraś, Maciej Ogrodniczuk, Michał Rudolf, Piotr Rybak, Karolina Saputa, Joanna Wołoszyn, Marcin Oleksy, Bartłomiej Koptyra, Teddy Ferdinan, Stanisław Woźniak, Maciej Piasecki, Paweł Walkowiak, Konrad Wojtasik, Arkadiusz Janz, Przemysław Kazienko, Julia Moska, Jan Kocoń

AI summary

Overview

Research area: Natural Language Processing — instruction dataset design for fine-tuning large language models, with a focus on Polish language and cultural adaptation.

Technical level: Intermediate. The paper assumes familiarity with supervised fine-tuning (SFT), preference alignment, and instruction-following evaluation, but explains its typology and design decisions in plain terms.

Scope: A description of the instruction corpus (PLLuMIC) built for the PLLuM (Polish Large Language Model) project, together with a 1,278-instruction public sample and experimental results on how instruction type affects model adaptation.

What This Paper Is About

Developers of both proprietary and open-weight LLMs rarely publish or adequately document the instruction data used to fine-tune their models, which makes capabilities hard to replicate and instructional design hard to learn from. The PLLuM project, funded by the Polish Ministry of Digital Affairs in 2024, needed an original instruction corpus to give a family of 8B to 70B parameter models their basic interactive capabilities in Polish. This paper documents that corpus, presents a functional typology of its organic, converted, and synthetic instructions, and reports what the authors learned about how these different instruction sources affect a model's linguistic adaptation.

Key Contributions

  1. A functional typology of instruction sources. The paper defines and contrasts three categories of instructions — organic (authored by humans, including experts, trained annotators, and collected human prompts), converted (derived automatically from annotated corpora, knowledge sources, dictionaries and ontologies), and synthetic (distilled from existing LLMs) — and discusses the advantages and limits of each.

  2. A documented instruction corpus. PLLuMIC combines 38,106 training organic instructions (49.12% of the training mix), 33,789 converted instructions (43.56%) and 5,679 synthetic instructions (7.32%), with total quantities of 47,295 organic, 33,789 converted and 5,679 synthetic instructions. Other languages covered include Ukrainian, Lithuanian, Russian and Belarussian.

  3. A public, representative subset. The first release contains 1,278 human-authored instructions spanning 12 types, 126 subtypes and 34 topics, available at https://huggingface.co/datasets/pelcra/PLLuMIC, with a synthetic extension (PLLuMIC-syn-ext) planned separately.

  4. Empirical evidence on instruction source effects. Evaluation on the Polish Linguistic and Cultural Competency Benchmark (PLCC), the LLMzSzŁ Polish exam benchmark, and a red-teaming suite shows that fine-tuning on PLLuMIC only helps models that have first undergone continual pre-training on the target language.

Main Findings

  • Continual pre-training is a prerequisite for instruction fine-tuning to help. Across all four architectures studied, continually pre-trained models outperformed their base counterparts. Fine-tuning on PLLuMIC alone degraded performance for models that had not been continually pre-trained: Mistral-Nemo-Instruct-2407 scored 23.00 on PLCC versus 22.33 for Mistral-Nemo-2407+PLLuMIC, Llama-3.1-8B-Instruct 22.67 versus Llama-3.1-8B+PLLuMIC 24.67, and Llama-3.1-70B-Instruct 47.83 versus Llama-3.1-70B+PLLuMIC 38.67.

  • Large PLCC gains follow continued pre-training plus PLLuMIC. PLLuM-12B-nc-instruct reached 56.33 and PLLuM-12B-nc-chat 59.50; PLLuM-8x7B-nc-instruct 67.17 and PLLuM-8x7B-nc-chat 68.17; Llama-PLLuM-8B-instruct 58.00 and Llama-PLLuM-8B-chat 60.67; Llama-PLLuM-70B-instruct 65.17 and Llama-PLLuM-70B-chat 66.33. Reference points on the same benchmark include Mixtral-8x7B-Instruct-v0.1 at 35.33, Qwen-Max at 50.83, GPT-4 at 59.50, Grok-2-1212 at 66.00, DeepSeek-v3 at 69.17, DeepSeek-R1 at 76.00 and O1-2024-12-17 at 89.17.

  • Alignment on human preferences adds a small PLCC gain. Aligned models scored slightly higher than their instruction-fine-tuned predecessors, which the authors partly attribute to aligned models generating longer responses and therefore more often meeting IFEval-style inclusion criteria.

  • Organic instructions fix specific, fine-grained language conventions. A subset of fewer than 100 high-quality e-mail writing instructions was enough to imprint Polish prescriptive conventions, such as lowercasing the second line after an addressative form followed by a comma, and avoiding a comma between complementary closings and newline signatures. These differ from English e-mail conventions.

  • Negative linguistic transfer is observable. Synthetic Polish e-mails produced by strong English-centric models translated formulaic English openings, and alignment on a preference dataset generated largely by English-dominant models reintroduced grammatical, lexical and stylistic inconsistencies that SFT had already fixed.

  • ORPO was the most effective alignment method tested. Odds Ratio Preference Optimization outperformed KTO, DPO and PPO, and applying ORPO after SFT still yielded superior results to alternative approaches.

  • Safety improved markedly after alignment. Red-teaming on 18,656 harmful prompts and 9,724 non-harmful samples, covering 14 Llama-Guard hazard categories and 10 attack styles inspired by the Rainbow Teaming framework, showed attack success rates of 1.03 for PLLuM-12B-nc-chat, 0.78 for PLLuM-8x7B-nc-chat, 0.76 for Llama-PLLuM-8B-chat and 0.79 for Llama-PLLuM-70B-chat, compared with 70.63 to 78.60 for the corresponding instruct models. False-refusal rates, however, were higher for the chat models than for the instruct models.

  • A safety-helpfulness trade-off appeared. Aligned models showed a tendency to refuse non-adversarial prompts, and verbosity increased even where a concise answer would suffice.

  • Converted instructions must be capped. Automatically converted instructions are highly repetitive, so a default maximum of 1,000 instructions was imposed per source resource.

  • General knowledge is competitive. On LLMzSzŁ, Llama-PLLuM-70B-chat scored 64.42, only 2.71 points below Llama-3.3-70B-Instruct at 67.13, which is reported to have been fine-tuned on millions of manually crafted instructions. Other scores: Llama-PLLuM-8B-chat 47.68, PLLuM-12B-nc-chat 53.40, PLLuM-8x7B-nc-chat 60.52, Meta-Llama-3.1-8B-Instruct 47.41, Mixtral-8x7B-Instruct-v0.1 49.46 and Bielik-11B-v2.1-Instruct 57.52.

Methodology in Plain English

The team built the corpus from three directions at once. First, professional human annotators wrote or adapted instructions from scratch, filling a high-level typology whose largest categories are Knowledge (QA) at 43% and Generation at 25% of the organic component, followed by Extraction and Programming at 6% each, Conversational 4%, and NLP, Adversarial, Visualization and Data manipulation at 3% each, Chain of Thought 2%, and Translation and Identity 1% each. Open datasets such as CREAK (3,591 samples), ECQA (1,033) and QED (1,855) were adapted early on, though the authors found many samples low-quality, simplistic or erroneous. Annotation shifted from single-turn prompt-response pairs to multi-turn dialogues, producing a subset of over 3,500 dialogues averaging approximately 12 turns each.

Second, synthetic instructions were generated through multi-step pipelines with minimal human supervision using locally hosted LLMs. Knowledge distillation moved from hand-written subject prompts to a meta-prompt that generated a question, then a second meta-prompt that generated the answer, all with Mixtral8x22b-instruct. Retrieval Augmented Generation instructions were built from Polish government documents in the gov.pl domain, with regular, adversarial and unrelated questions; the top 5 documents per fragment were retrieved using the bge-m3 retriever and bge-reranker-v2-m3 reranker, answers were generated by Llama-3.3-70B as a strong reference and preference answers by Llama-3.1-8B as a weaker model. RAG instructions were capped at 5,000 for SFT and 5,000 preference pairs, with a final training mix of 80% regular, 14% adversarial and 6% unrelated questions. Context-injected NLP tasks used text samples injected into system prompts specifying structured outputs such as JSON, CSV and XML.

Third, converted instructions were produced automatically from annotated corpora, treebanks, named-entity datasets, machine-readable dictionaries and ontologies, using hand-written question-and-answer templates; a prompt format was chosen at random when mapping an example to a single-turn instruction, only train splits were used for training, and validation splits were optionally used for internal evaluation.

For the language adaptation experiments, base models were continually pre-trained on a corpus of approximately 150 billion tokens, annealed on a subset, then fine-tuned for 3 epochs with the AdamW optimizer (weight decay 0.1), a learning rate of 1e-5, a cosine scheduler with 1% warmup, and a cumulative batch size of 128, with a maximum sequence length of 16,384 tokens; loss was computed only on the response turn. Training used a multi-node configuration on NVIDIA H100 nodes at the Wrocław Centre for Networking and Supercomputing with DeepSpeed ZeRO Stage 3. Alignment used over 40,000 manually annotated instructions collected through rating-based, ranking-based (four responses ranked) and dialog-based annotation, with over 50 different annotators. Evaluation used PLCC (600 questions covering Polish history, geography, culture, tradition, art, entertainment, grammar and vocabulary, scored with an IFEval scheme), LLMzSzŁ, and the red-teaming suite.

Why This Matters

Impact on research. Most fine-tuning data behind released LLMs is withheld or thinly documented; an overview of open-weight models in the paper shows the vast majority come with no instructions and little documentation. This paper gives a documented, replicable account of how a national-scale instruction corpus was designed, which is rare and useful for teams planning comparable resources in other languages.

Real-world applications (as grounded in the paper's domains):

  • Public administration question answering over gov.pl documents, covering identity documents, business activity, taxes and residence registration.
  • RAG systems that must handle adversarial and unrelated questions gracefully, using the 14% adversarial and 6% unrelated proportions tested here.
  • Polish-language professional writing assistance, such as e-mail composition following prescriptive punctuation and formatting conventions.
  • Safe deployment in culturally specific contexts, where red-teaming results quantify attack success and false-refusal rates.

Industry relevance. The finding that instruction fine-tuning only works after sufficient continual pre-training has direct implications for any organization trying to adapt an English-dominant open-weight model to a lower- or mid-resourced language: skipping the language priming step can actively degrade performance. The paper also documents legal constraints on synthetic data generation from licensed LLMs and the risk of model degradation or collapse from uncontrolled recursive distillation, both of which are practical planning concerns.

Future Directions

  1. Release and evaluate the synthetic extension. The authors state they plan to release PLLuMIC-syn-ext separately, which would allow direct comparison of the public organic sample against synthetic data.

  2. Quantify instruction-type effects more precisely. The paper reports ablation experiments appended to the appendix addressing the relationship between continual pre-training and instruction fine-tuning, but the relative contribution of organic versus converted versus synthetic data remains an open question for other architectures and languages.

  3. Reduce negative linguistic transfer during alignment. Alignment on preference data generated by English-dominant models reintroduced errors that SFT had fixed, raising the question of how to build preference datasets that preserve target-language conventions.

  4. Rebalance safety and helpfulness. Aligned models showed higher false-refusal rates than instruct models, so methods that retain the low attack success rates while avoiding over-refusal on benign prompts remain to be developed.

Target Audience

This paper is most useful to researchers and engineers building instruction datasets or fine-tuning LLMs for languages other than English, particularly low- and mid-resourced languages. It also suits project leads planning national or domain-specific language model programmes, annotation managers who need a workable typology and quality-control model, and NLP researchers studying language transfer, alignment methods, or the relative value of human versus synthetic instruction data. Readers looking for a benchmark-only or architecture-focused paper will find less here; the emphasis is on data design and documented practice.

Authors’ abstract

This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project. We present a functional typology of the organic, converted, and synthetic instructions used in PLLuM and share some observations about the implications of using human-authored versus synthetic instruction datasets in the linguistic adaptation of base LLMs. Additionally, we release the first representative subset of the PLLuM instruction corpus (PLLuMIC), which we believe to be useful in guiding and planning the development of similar datasets for other LLMs.

Read the original paper