Skip to content
AI.info

Research

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

Overview Research area: Natural Language Processing / Large Language Model evaluation, with a focus on low-resource and under-served languages (Slovak). Technical level: Intermediate. The paper is rea

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
arXiv
2610.09152
Published
2026-10-06
Authors
Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, Daniel Hládek, Michal Gregor, Martin Tamajka, Marián Šimko

AI summary

Overview

  • Research area: Natural Language Processing / Large Language Model evaluation, with a focus on low-resource and under-served languages (Slovak).
  • Technical level: Intermediate. The paper is readable without deep ML background, but it assumes familiarity with benchmarks, leaderboards, and terms like "zero-shot," "instruction tuning," and "test-time reasoning."
  • Scope: The paper introduces sk-bench, a native-first Slovak evaluation suite of 30 datasets (33 scored task variants) across ten skill categories, used to evaluate 55 open- and closed-weights models under a single harness.

What This Paper Is About

Multilingual LLM benchmarks either leave Slovak out entirely or include it only through machine translation, and no unified harness existed for comparing open and proprietary models on Slovak generative tasks. The authors build sk-bench, a Slovak-first benchmark that prioritizes natively authored data over translation, and use it to answer three questions: how much data provenance (native vs. translated vs. LLM-generated) changes model rankings, what adapting a model to Slovak costs, and whether test-time reasoning transfers to Slovak.

Key Contributions

  1. sk-bench itself: a native-first benchmark for generative-LLM evaluation in Slovak, with 30 tasks across ten skill categories, run under one harness against 55 contemporary open- and closed-weights models, with a unified Slovak leaderboard.
  2. Eleven new Slovak evaluation resources: five built from scratch, four curated from public institutional sources where no dataset previously existed, and two existing expert resources first packaged for generative-LLM evaluation. Two of these (IFEval-SK with a Slovak instruction-checker registry, and the paired Chiby/SKJ1 grammaticality probe) could not have been built by translation.
  3. A reproducible harness and open release: a generate-and-match pipeline scoring open-weights and API models identically, with prompts, run scripts, and per-task breakdowns released publicly.
  4. Four transferable design lessons (L1–L4) for benchmarks in other under-resourced languages, each backed by a Slovak-specific empirical result.

Main Findings

  • Proprietary APIs lead the Slovak leaderboard: Gemini 3.1 Pro tops the overall score at 72.6, ahead of GPT-5.5 (71.0) and Claude Opus 4.6 (68.1). The strongest open-weights model is gpt-oss 120B at 60.0, so the best open model trails proprietary APIs by 12.6 points.
  • Scale is a weak predictor of Slovak ability: open-weights scores plateau past roughly 20B parameters, and no larger open model up to the 1.1T Kimi K2.6 beats gpt-oss 120B (60.0). Mid-sized models stay competitive: Qwen3.5 27B (55.3), Gemma 4 31B (54.3), and gpt-oss 20B (54.1).
  • Translation is a safe ranking proxy for closed-form skills: native vs. translated NLI rankings correlate at ρ = 0.99 (95% bootstrap CI [0.97, 1.00], identical top-10); native vs. translated multiple choice correlates at ρ = 0.98 (CI [0.94, 0.99]), but the top-10 lists share only 7 of 10 models, suggesting translated MMLU separates the strongest models less reliably.
  • Question provenance reorders QA rankings: human-authored skLEP-QA and LLM-generated MultiWikiQA-SK agree only at ρ = 0.72 (CI [0.57, 0.83], top-10 overlap 6/10), with the disagreement largest below the frontier (ρ = 0.17 at <2B and 0.43 at 2–9B, rising to 0.74 at 9–32B).
  • Slovak adaptation is costly: continued Slovak pretraining for Qwen3-14B lowers the overall score from 47.4 to 33.5. It raises the grammar group by +6.8 but degrades knowledge (-23.7), reading (-25.2), and classification (-30.6); instruction following falls only -6.4.
  • A small instruction set largely repairs the loss: for the Qwen lineage, Alpaca-style Slovak instruction tuning restores the score to 43.8, about three-quarters of the 13.9-point deficit, while still 3.6 points below the original Qwen3-14B.
  • Instruction repair is base-dependent and rises with base capability: Mistral-SK (19.0 overall) loses 4.0 points, TildeOpen-SK (31.9) is essentially flat (+0.8), and Qwen3-SK (33.5) gains +10.3. Instruction following improves for two bases (+13.2, +11.5) and is unchanged for the third.
  • Test-time reasoning transfers to Slovak at scale: nine models with an optional thinking mode were evaluated under an 8,192-token cap. Every model of 9B or more gains +8.5 to +12.5 overall, concentrated in instruction following (IFEval-SK improves for every model, by +7 to +54 points, and by at least +21 above 9B). Free-form generation is flat or slightly negative (≤ +0.2).
  • Reasoning backfires below 9B: Qwen3.5 0.8B and 2B regress by -16.8 and -9.7 overall, a format failure where models exhaust the budget inside the trace without emitting a parseable answer.
  • Prompt language barely matters: switching every template to English changes the overall score by only +0.4 points; grammar & morphology is the only sizable shift (-2.6, where Slovak helps).
  • Few-shot examples give modest gains: 3-shot raises the overall score across the 52 models with few-shot results by +2.0 (Slovak) and +1.6 (English), helping classification and reading & QA but slightly lowering knowledge and reasoning.
  • No Slovak length cliff up to 16k: on OneRuler-SK, Slovak retrieval declines only mildly from 4k to 16k (58.1 to 54.4) and tracks English throughout.
  • Summarisation rankings are metric-dependent: under single-reference ROUGE-L, Llama 3.3 70B (21.4) tops the generation group ahead of GPT-5.5 (15.5) and Gemini 3.1 Pro (14.8), yet under chrF Llama falls to 13th (28.6, behind both at 29.0), and under METEOR the best summariser is Claude Opus 4.6. The authors treat these as relative rankings, not absolute measures.

Methodology in Plain English

The authors assembled a Slovak benchmark by preferring natively authored data over translations, and only admitting translated datasets when no native equivalent existed for a skill. Selection followed four criteria: native data preferred, redistribution-friendly licence, at least 100 test examples, and automatic reproducible scoring.

The suite combines three kinds of data: native Slovak data testing language-specific competence (grammar, morphology, sentiment, toxicity, NLI, professional and school exams), Slovak domain data with professionally validated labels, and translated or generated proxies for skills with no native benchmark. Where possible, the same skill is measured from two differently sourced datasets — a "paired provenance contrast" — so provenance can be varied while the skill is held fixed.

All models were run through one pipeline built on LightEval: open weights served with vLLM, closed models queried through LiteLLM. Every task is scored by generating text up to a per-task token budget (32 tokens for classification and multi-choice, 512 for QA and summarisation, 4096 for IFEval, 8192 in reasoning mode), then matching strings against the gold answer. Default settings were zero-shot, greedy decoding (temperature=0, fixed seed=1234). Hand-written Slovak instruction templates were used for the main results, with parallel English templates as a robustness check, plus a 3-shot variant of every task.

The main sk-bench score is an unweighted mean of six reporting-group averages, with task macro-averaging inside each group so that large tasks such as the 13,062-item Okapi-MMLU-SK cannot dominate against the 500-item SlovakCOPA. To create Slovak instruction-tuned counterparts, three base models (Mistral-SK-7B, Qwen3-14B-sk, TildeOpen-30b-64k) were fine-tuned on the Slovak split of TaCo Alpaca data.

Why This Matters

Impact on research: The paper provides a template for how to build evaluation suites in under-resourced languages, and it empirically tests an assumption many multilingual benchmarks make implicitly — that machine-translated data is a safe stand-in for native data. Its finding that translation is safe for closed-form skills (ρ ≥ 0.98) but not for QA (ρ = 0.72) gives benchmark builders a concrete rule about where to invest native annotation effort. The released per-task breakdowns also let others re-aggregate under different weightings.

Real-world applications:

  • Slovak-language customer support and chatbot deployment, where practitioners need identically prompted scores to choose between open-weights and API models.
  • Public administration and citizen-facing services requiring reliable Slovak text understanding.
  • Slovak education and admissions, since the benchmark includes natively sourced exam sets (driving law, admission tests, school subjects) with professionally validated gold labels.
  • Content moderation and sentiment analysis for Slovak news and social media, drawing on the SentiSK and ToxicSK tasks.
  • Slovak summarisation pipelines for news, tested on SlovakSum and SME-Sum.

Industry relevance: The open-versus-closed comparison under a single harness (in the spirit of MEGAVERSE) directly supports deployment decisions. The adaptation findings matter to any organization that has continued-pretrained or fine-tuned a model on a target language: the Qwen3-14B case shows a 13.9-point overall regression from Slovak pretraining, and that a small instruction set recovers about three-quarters of it — a cheap repair step worth planning for. The reasoning results give a capacity threshold (9B) above which thinking mode pays off in Slovak, and below which it actively harms scores.

Future Directions

  1. Item-matched translation controls. The authors note that the two translation-contrast rows pair differently sourced datasets covering the same skill, mixing content differences with provenance. An item-matched control would isolate translation per se, and they say it is constructible from released data since Belebele-SK and Global-MMLU-Lite are parallel across Slovak and English.
  2. Broadening coverage. sk-bench currently covers only textual single-turn evaluation in standard written Slovak, and does not cover dialogue, multi-turn agentic tool use, multimodality, code generation, or dialectal/colloquial Slovak. Each is an open extension.
  3. Test-set contamination measurement. The authors have not measured or certified the absence of sk-bench test items from evaluated models' training corpora, and they distinguish three leakage layers (benchmark-format, answer-key, source-document). Quantifying this remains open, particularly for closed models with undisclosed training data.
  4. Aggregation sensitivity. The score weights all reporting groups equally with task macro-averaging; the authors acknowledge an example-weighted score would yield different rankings and invite re-aggregation from the released per-task breakdowns.

Target Audience

Researchers and engineers working on multilingual or low-resource NLP; benchmark designers deciding how much native data to collect versus how much translation to accept; practitioners selecting an LLM for a Slovak-language deployment; and teams that have adapted or plan to adapt an instruction-tuned model to a new language and want to know the expected capability cost and repair path.

Authors’ abstract

Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($ρ\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($ρ=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench

Read the original paper