Research
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
Overview Research area: Applied natural language processing and LLM post-training for enterprise deployment, with focus on Russian-language instruction following and function calling. Technical level:

- arXiv
- 2609.01572
- Published
- 2026-09-01
- Authors
- Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov, Danil Taranets, Daniil Dryabin, Mikhail Gashkov, Viktor Zelenkovskiy, Aleksandr Fida, Gleb Alektorov, Nikita Gulyakov, Arthur Babkin, Aleksandr Medvedev, Pavel Gein, Anatolii Potapov
AI summary
Overview
Research area: Applied natural language processing and LLM post-training for enterprise deployment, with focus on Russian-language instruction following and function calling.
Technical level: Intermediate to Advanced. The paper assumes familiarity with supervised fine-tuning (SFT), reinforcement learning from verifiable rewards, GRPO, reward models, and weight-space merging; the surrounding production-benchmarking methodology is accessible to a general ML audience.
Scope in one sentence: The paper describes how one company consolidated traffic from over 200 internal applications onto a single self-hosted model by diagnosing three production failure axes, training a separate GRPO expert per axis, and merging the experts via two-stage SLERP.
What This Paper Is About
Enterprises bound by data-residency rules must run LLMs locally, but as newer open-weight models arrive without older ones being retired, a single finite GPU pool gets split across a growing fleet of models and the effective price per token rises. The authors' goal is to migrate the traffic of hundreds of internal applications onto one Qwen3-32B-based model by closing the specific quality gaps that block migration: instruction following, function calling, and alignment to the internal request mix, without regressing general capabilities.
Key Contributions
-
A methodology for building internal benchmarks from production traffic, scored by deterministic verifiers or calibrated LLM judges, and validated against human annotators. It includes a template-aware sampler (diversity 0.953 while keeping Jensen-Shannon distances close to production) and a task-specific judging pipeline that lifts Cohen's κ from 0.63 to 0.88 on reference-based tasks and from 0.57 to 0.72 on open-ended content generation.
-
A modular post-training recipe that trains a separate GRPO expert per weak axis from one shared SFT checkpoint and combines them by two-stage sequential SLERP merging, rather than jointly optimizing all objectives.
-
Documentation of three reward-hacking failure modes that make joint multi-objective training fragile, each with a domain-specific fix: semantic collapse (instruction following), over-calling (function calling), and verbosity hacking (general/dialogue).
-
An open-weight checkpoint trained with the same recipe but without internal data (https://huggingface.co/t-tech/T-pro-it-2.1), which scores close to the deployed version on public benchmarks, supporting the claim that the recipe, not proprietary data, drives the gains.
Main Findings
-
Single model absorbs the fleet: Six months after rollout the adapted model absorbs 50% of platform traffic, 116M requests per month, from over 200 internal applications, at a fraction of the serving cost.
-
Shared SFT is enough; joint GRPO is not: A single shared SFT stage matches per-domain SFT experts across the in-house Arena (92.45 vs 91.96), ruIFEval (66.00 vs 62.18), and BFCL English (0.787 vs 0.798). At the GRPO stage, however, joint multi-objective training fails to hold all domains at once.
-
Domains do not transfer through the reward: Single-domain GRPO gains stay confined to their own axis. The general expert lifts the arena scores to 95.26 and 70.73 but leaves IFEval and tool-calling near the SFT baseline; the IF and FC experts behave the same way.
-
Joint GRPO vs. merging at 32B: The merge gives in-house Arena 93.87, ruIFEval 69.57, BFCL English 0.799, Russian 65.96, and in-house BFCL 72.27. Joint GRPO from the SFT checkpoint falls to ruIFEval 0.770, BFCLv3 English 70.38 and Russian 60.88; the general-GRPO warm start reaches parity only with a 1.7× larger budget.
-
Merge order matters: Because SLERP is non-associative, composing the two verifiable-reward experts first and then the general expert, (IF + FC) + Gen., outperformed the other orderings with in-house Arena 93.37, ruIFEval 68.99, BFCL 0.798, Ru-Hard 72.19, and in-house 65.73.
-
In-house reward-model adaptation backfires: Mixing in-house preference pairs into reward-model training did not help (in-house Arena 68.58 vs 70.73 for the general reward model) and produced longer responses, averaging 362 tokens versus 286.
-
Quality versus a ~7× larger by total parameters baseline: In non-reasoning mode the final model reaches in-house Arena 69.57 vs 65.83 for Qwen3-235B-A22B-Instruct-2507, in-house BFCL 0.79 vs 0.77, ruBFCLv3 65.96 vs 64.42, and AceBench 73.50 vs 70.20. The abstract states the recipe surpasses the ~7× larger baseline on the in-house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function-calling with 0.79 to 0.77.
-
Gaps narrow where backbone scale dominates: SmartSearch F1 improves from 0.478 to 0.557, surpassing both 32B-scale thinking-mode baselines, and ruWildChat rises from 52.0 to 80.7, within 4.4 points of the ~7× larger model. ruMultiChallenge remains lower (34.1 vs 46.2 at 37.60 vs 40.97 on τ²-bench).
-
Failure taxonomy from production: Over a human-reviewed sample of n=2,500 responses (κ=0.62), classification is the largest single category at 36.0%, but combined instruction-following failures reach 37.9% (21.0% formatting plus 16.9% non-format). Tool-equipped requests are ~12% of traffic.
-
Task-specific judging beats uniform judging: A uniform side-by-side judge agreed with expert annotators at κ=0.62. A task classifier matched human consensus in 90.6–99.6% of cases; classification and information extraction (~63.2% of traffic) were scored against gold answers from Kimi-K2.5, accepted unmodified at 97.7% and 85.2%. Summarization was best judged with a checklist as contextual guidance (κ=0.68), content generation with per-criterion grading plus an overall verdict (κ=0.79).
-
Deployment profile: 116M requests per month, 45 average and 110 peak requests per second, single-GPU FP8 replicas behind vLLM, 16 to 48 pods, 95th-percentile latency of 3.2 s and time-to-first-token of 0.3 s. Per-token cost falls by 2.8 to 3.9× on input and output versus the ~7× larger baseline, and up to 4 to 9× for services that previously ran the largest platform models.
-
Each reward needed a fix: The instruction-following verifier led to minimal, semantically empty completions, corrected with a prompt-specific reward-model quality penalty; the Tool-N1 exact-match reward encouraged always calling, corrected by injecting synthetic irrelevance data; the general expert used a multiplicative length penalty against a Qwen3-235B-A22B-Instruct-2507 baseline plus an increased KL coefficient.
Methodology in Plain English
The authors start from the traffic already flowing through their internal LLM platform. They sample roughly 100k monthly queries with a sampler that masks variable tokens, groups near-identical templated prompts via locality-sensitive hashing, and selects within each template by greedy max-min, allocating budget as the square root of template count. Each sampled request is routed by an LLM task classifier to a task-specific scoring scheme, judged either by deterministic verifiers against gold answers or by an LLM judge whose agreement with human annotators is measured with Cohen's κ. Human annotators label failure types and judge-validation pairs, with three annotators per pair and majority vote.
On the training side they begin with Qwen3-32B and an adapted Cyrillic-dense tokenizer, run one combined SFT stage mixing in-house production, general-domain, instruction-following, and function-calling data, and then fork that checkpoint into three independent GRPO runs, one per axis. Instead of trying to balance all three rewards in one run, they merge the resulting expert checkpoints in weight space using two-stage sequential SLERP, choosing the merge order empirically. Instruction-following data comes from a Russian adaptation of AutoIF, expanding 54 hand-written seed constraints into 43K verified constraints and 26K training examples. Function-calling data is generated natively in each language, 1.2M English and 300K Russian samples, using a planner-and-simulator dialogue pipeline so the training targets reflect realistic tool use. The general expert mixes roughly 80% general-domain Russian instruction data with 20% in-house samples and uses Qwen3-235B-A22B-Instruct-2507 as the regeneration teacher. The model operates only in non-reasoning mode because of production latency and cost limits.
Why This Matters
Impact on research: The paper argues that in enterprise post-training the reward signal, not the data volume, is the bottleneck, and shows concretely that joint multi-objective GRPO creates cross-domain interference while per-axis experts plus weight-space merging do not. It also documents three reward-hacking modes and the domain-specific data or reward corrections each requires, which is reusable evidence for anyone designing multi-objective RL post-training.
Real-world applications:
- Internal enterprise assistants and support tools that must run on self-hosted infrastructure under data-residency rules.
- Russian-language tool-using agents, where the paper locates the main failure not in intent understanding but in Russian-language argument filling and Russian-documented tools.
- Cost consolidation programs that need to retire a fragmented fleet of models and standardize on one checkpoint per GPU pool.
- Benchmark construction from live traffic for teams whose long-tail applications cannot accumulate enough traffic for reliable A/B testing.
Industry relevance: The deployment numbers, 116M requests per month across over 200 services on single-GPU FP8 replicas with 2.8 to 3.9× per-token cost reduction, show a concrete operating point for a 32B dense non-reasoning model. The authors note the few rollbacks came from teams requiring frontier-scale agentic capabilities beyond a 32B dense model.
Future Directions
- Validating the recipe outside Russian and English, and outside a single self-hosted corporate deployment, since all quantitative claims are stated as validated for those two languages only.
- Testing the recipe on model families other than Qwen3, since all experiments here use only that family.
- Recalibrating LLM judges against human annotations when the recipe is reused for another deployment, language, or benchmark distribution; the paper reports a sensitivity study over four judges in its appendix.
- Investigating the residual gaps on tasks where backbone scale dominates, namely ruMultiChallenge for long-context memory and self-coherence and SmartSearch for open-ended retrieval and synthesis, which the recipe narrows but does not close.
Target Audience
ML engineers and applied researchers building or consolidating self-hosted LLM serving for enterprises, especially those working in Russian or other non-English enterprise contexts. It is also relevant to post-training practitioners interested in multi-objective RL, reward hacking, and model merging, and to evaluation engineers who need to construct representative internal benchmarks from production traffic and calibrate LLM judges against human annotators.
Authors’ abstract
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimising all objectives jointly, which introduces cross-domain reward interference, we train a separate GRPO expert per axis and merge them via two-stage SLERP. Each expert's reward exposes a distinct failure mode, namely semantic collapse, over-calling, and verbosity hacking, each requiring a domain-specific fix. In non-reasoning mode the recipe surpasses a ${\sim}7\times$ larger by total parameters baseline on the in-house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function-calling with 0.79 to 0.77, while lifting general dialogue benchmarks. The model absorbs 50% of platform traffic, 116M requests per month, at a fraction of the serving cost.