Skip to content
AI.info

Research

Miles v0.1: Production-Level Post-Training

Overview Research area: Machine learning systems — post-training infrastructure for large language models (reinforcement learning, supervised fine-tuning, distillation, and agentic rollouts at frontie

Miles v0.1: Production-Level Post-Training
arXiv
2609.08368
Published
2026-09-08
Authors
RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng

AI summary

Overview

  • Research area: Machine learning systems — post-training infrastructure for large language models (reinforcement learning, supervised fine-tuning, distillation, and agentic rollouts at frontier scale).
  • Technical level: Advanced. The paper assumes familiarity with RL training loops, mixture-of-experts (MoE) routing, low-precision number formats, and distributed training frameworks (Megatron-LM, FSDP, SGLang).
  • Scope: A system report describing the architecture, components, configuration knobs, and measured behavior of Miles v0.1, a full-stack post-training system, closing with one end-to-end agentic RL case study.

What This Paper Is About

Post-training is what turns a pretrained language model into a useful one, but at frontier scale the training loop has outgrown a simple generate-then-update cycle: rollouts are multi-turn, use tools, act in external environments, and are often produced by trillion-parameter MoE models. The core problem is that these workloads make it hard to keep expensive hardware busy — latency-sensitive rollout generation and throughput-oriented training create idle bubbles — while numerical differences between the serving engines and the trainer can silently corrupt the training objective. Miles v0.1 is presented as a full-stack, production-ready system that addresses both problems, organized around the principle that every component should be verified, clean, and customizable.

Key Contributions

  1. A three-stage RL loop with a fully asynchronous mode. Miles cycles through rollout (SGLang engines), training (Megatron-LM or PyTorch FSDP), and weight update. A fully asynchronous mode lets generation and training progress concurrently on separate (disaggregated) GPU pools, with a bounded data buffer decoupling the two stages.

  2. Fidelity mechanisms for agentic rollouts. Cache-preserving request routing ("affinity"), a Token-In-Token-Out (TITO) session server that keeps the trainer's view of a trajectory token-exact across turns, and rollout routing replay (R3) that replays each token's MoE expert assignments during training instead of recomputing them.

  3. A shared low-precision contract between rollout and training. FP8 blockwise, MXFP8, and NVFP4 recipes are built as end-to-end "precision contracts" spanning four stages — checkpoint conversion, trainer forward pass, SGLang rollout, and live weight export — so both sides quantize the same weights the same way.

  4. Extensibility beyond core RL. LoRA RL, on-policy distillation, supervised fine-tuning, true-on-policy rollout-training alignment, and extension of the same architecture to diffusion models, plus nested plug-in layers for agentic environments and a multi-vendor hardware coverage statement.

Main Findings

  • Asynchrony removes idle time from turn-taking. Under a synchronous schedule, a batch completes only after its slowest trajectory, leaving most GPUs idle; under Miles's asynchronous schedule the rollout engines generate continuously and the trainer pulls whichever groups have already finished. Miles refuses to start a fully asynchronous run when the trainer and rollout engines share GPUs (colocated placement).

  • Sample-granularity replacement is the default under fully asynchronous rollout. Group granularity waits for every trajectory in a group before starting a replacement; sample granularity frees each finished trajectory's place immediately, keeping the number of in-flight trajectories near the limit even when trajectory lengths differ by an order of magnitude.

  • Affinity plus least-loaded placement achieves a 96% prefix-cache hit rate. Miles binds a session to the engine holding its KV cache (and to a data-parallel rank when DP attention is enabled), and the session server sends each new session to the engine with the fewest active requests. A trajectory reaching the router without a routing key raises an error rather than silently falling back to load-based routing.

  • Staleness is defined pessimistically. A group's staleness is the current trainer weight version minus the oldest weight version appearing anywhere in the group, so a group is never treated as fresher than its oldest token. Staleness is checked only when the trainer collects a group; the other two drop conditions (generation gave up, or a user filter rejected it) are checked on arrival.

  • Evaluation must be tied to a policy version. Miles offers three asynchronous evaluation modes: shared rollout engines (default, pauses generation), a dedicated evaluation fleet (loads a checkpoint snapshot), and an external backend (checkpoint directory). A dedicated fleet verifies weight delivery before evaluating; Miles checks that the weight version averaged across samples equals the step the score was recorded against, and that the fraction of requests served by a mixed set of versions is zero. Evaluation failures never stop training — they are recorded as skipped with a logged reason.

  • Token exactness depends on configurable replay comparison. Miles supports linear and branching session extension rules and three built-in comparison settings ranging from strictest (comparing only the fields a chat template reads) to loosest (comparing only role and visible text, ignoring tool calls entirely). The loosest setting can merge genuinely different histories, and the paper notes Miles does not reconcile tool-call identifiers when histories collapse this way, so the mismatch remains silent.

  • TITO registration covers specific model families. Registrations span the Qwen3, GLM, Nemotron, Kimi, MiniMax, DeepSeek, and Inkling lines; checkpoints outside those lines fall back to a generic handler with a warning. Two checks guard each registration: a CPU check for append-only token sequences, and a GPU check against a live model. The session server does not yet carry image or video inputs, so vision-language models drive the SGLang engine directly.

  • R3 has a real memory cost. Each routing tensor holds (tokens − 1) × layers × k 32-bit integers; for a 32K-token sequence over 60 layers at k = 8, that is roughly 60 MB per trajectory. For a 32K-token sequence over 60 layers at k=8, that tensor occupies roughly 60 MB per trajectory. R3 is a per-recipe choice: it does nothing for dense models, may have limited effect in asynchronous RL, and the GLM-5.2 reference run leaves it off.

  • Three low-precision formats have end-to-end recipes. FP8 blockwise (128×128 blocks, FP32 scales) is generally available and tested on Qwen3-4B, Qwen3-30B-A3B, and DeepSeek-V4. MXFP8 (1×32 blocks, UE8M0 scales) is Beta and tested on Qwen3-30B-A3B and DeepSeek-V3.2. NVFP4 (E2M1, 1×16 blocks, E4M3 and FP32 scales) is Beta and tested on Qwen3-30B-A3B. FP8 blockwise runs on NVIDIA Hopper and Blackwell and on AMD MI350X and MI355X; MXFP8 and NVFP4 require Blackwell; A100 has no FP8 arithmetic and runs BF16 only.

  • Format mixing is restricted. Rollout and training may either run the same format, or the trainer may stay in BF16 while rollout quantizes. NVFP4 requires every stage that touches the weights to quantize them, so pairing NVFP4 rollout with a BF16 trainer is unsupported. MXFP8 and NVFP4 hold a few tensors in BF16 as per-tensor exceptions, and overrides may be applied by layer range and tensor name at each of the four contract stages.

  • Two optional NVFP4 refinements differ in scope. "Dequantized backward" touches only the training side, running backward GEMMs in BF16 on operands dequantized from the NVFP4 values used in the forward pass. "Four-over-six" changes how each NVFP4 block is quantized and therefore belongs to the contract — it is enabled in both the Transformer Engine kernels the trainer uses and the FlashInfer kernels SGLang uses.

  • Precision verification is by log-probability comparison. On the configurations measured so far, reward curves track the BF16 baseline closely and rollout time is reduced significantly; the same quantization might behave differently on different models.

  • Two offloading mechanisms compose. Evicting the paused actor moves the entire training process off the GPU between steps; streaming the optimizer state keeps optimizer memory off the GPU during the step and fetches it back one bucket at a time. A run can do both.

  • The end-to-end case study reports 263 seconds. Fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, on 64 NVIDIA GB300 GPUs, with a median step time of 263 seconds over the first 30 measured steps.

Methodology in Plain English

The authors built a training system rather than proposing a new algorithm. They started from the design of slime and organized every stage of the RL loop around three properties: components should be verified (checked before they are trusted), clean (small, replaceable, with clear boundaries), and customizable (a user can swap a piece without rewriting the loop).

Concretely, they route all requests from one multi-turn episode to the same inference engine so the prefix stays in that engine's KV cache, and they let a session server — not the agent's harness — decide how messages become tokens, so the trainer later sees the exact tokens the model sampled. When rollout and training would otherwise take turns, they run them concurrently on separate GPU pools and place finished groups in a bounded buffer that also decides when a group is too stale to train on. To make low-precision training safe, they enforce a single quantization contract applied identically at checkpoint conversion, the trainer's forward pass, rollout, and the live weight export, then check the result by comparing log-probabilities assigned by SGLang and Megatron-LM to the same sampled tokens. For MoE models they optionally record which experts each token was routed to during rollout and replay those assignments during training. They then measure a real agentic RL run end to end.

The paper also states its own limits: some precision formats remain at an early stage, some weight-transfer paths cover only certain model families, and some measurements come from one configuration rather than many. The connectors and sandbox integrations named are described as experimental, and the report's later sections (weight synchronization details, LoRA RL, distillation, true-on-policy alignment, diffusion support, day-0 model support, multi-vendor hardware coverage, and the full case study) are referenced but not contained in the excerpt provided.

Why This Matters

Impact on research. The paper's central claim is that train-rollout mismatch is a silent failure mode, not a performance tuning issue: if the trainer's log-probabilities do not match what the policy actually sampled, the gradient is computed against a trajectory that never occurred, and the importance ratio drifts away from one. By treating token exactness, expert-routing replay, and a shared quantization contract as correctness requirements rather than optimizations, Miles presses on a problem that many RL-for-LLMs pipelines leave implicit. The fully asynchronous schedule also reframes the straggler problem as a system-design question rather than a batch-construction one.

Real-world applications:

  • Large-scale agentic RL for coding agents that run commands, edit files, and are graded by test suites, as in the terminal-use case study.
  • Reinforcement learning on trillion-parameter MoE models, where expert routing mismatch is a documented destabilizer (the paper cites Ma et al. for routing discrepancy destabilizing RL in MoE models and ending in catastrophic training collapse).
  • Low-precision, cost-conscious post-training on 4-bit and 8-bit formats, where a shared quantization contract is needed to keep rollout and training numerically consistent.
  • Enterprise post-training pipelines that need to plug in their own environments, sandboxes, task sets, and data selectors without modifying the training loop.

Industry relevance. The work is aimed explicitly at both researchers and enterprises, spanning two trainer backends, three weight-synchronization transports, four sandbox providers (AgentENV, Daytona, E2B, Modal), and multi-vendor hardware (NVIDIA Hopper, Blackwell, and AMD MI350X/MI355X for FP8 blockwise). The emphasis on observability metrics for the rollout-training buffer reflects a production concern: two stages advancing at independent rates can waste hardware without crashing or raising an error.

Future Directions

  • Vision-language support through the session server. The session server does not carry image or video inputs yet, so computer-use connectors record trajectories at the generate-function layer; once it does, they could attach at the agent function instead. The paper also notes the session server does not record screenshots yet.
  • Maturing the Beta precision recipes. MXFP8 and NVFP4 remain Beta and have been tested on a limited set of models; the authors state the same quantization might behave differently on different models. INT4 quantization-aware training and a "BF16 train, FP8 serve" mode are mentioned as further options but are not detailed in the excerpt.
  • Broadening weight-transfer coverage and multi-model verification. The paper states that some weight-transfer paths cover only certain model families and that some measurements come from one configuration rather than many.
  • Reducing or avoiding the cost of routing replay. R3's routing tensors are cheap to replay but expensive to carry (roughly 60 MB per trajectory for a 32K-token sequence over 60 layers at k = 8), and its benefit is limited in asynchronous RL and absent for dense models — leaving open how to keep its stability benefit without its overhead.

Target Audience

This paper is most useful to ML systems engineers and infrastructure teams building or operating large-scale post-training pipelines; RL researchers working on MoE models, agentic rollouts, or low-precision training who need to understand failure modes like train-rollout mismatch; and technical decision-makers evaluating a post-training stack for frontier-scale or enterprise deployment. Readers need working knowledge of distributed training, reinforcement learning objectives (the paper mentions GRPO among its targets), and low-precision number formats to follow the configuration-level detail.

Authors’ abstract

We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.

Read the original paper