Skip to content
AI.info

Research

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

Overview Research area: Distributed systems and systems for machine learning (cs.DC) — specifically, memory-efficient full-parameter fine-tuning of large language models across multiple GPUs that shar

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
arXiv
2609.34162
Published
2026-09-28
Authors
Ruijia Yang, Shiyuan Lin, Yulong Ao, Zhiyu Li, Yingli Zhao, Xianduo Li, Yonghua Lin, Zeyi Wen

AI summary

Overview

Research area: Distributed systems and systems for machine learning (cs.DC) — specifically, memory-efficient full-parameter fine-tuning of large language models across multiple GPUs that share one host machine.

Technical level: Advanced. The paper assumes familiarity with data parallelism, mixed-precision Adam training states, host–device transfer, and pipeline scheduling.

Scope: SlideDP is a synchronous data-parallel runtime that keeps one authoritative copy of master weights and optimizer states in CPU memory and streams layers through a bounded GPU window, coordinating parameter delivery, gradient aggregation, and layer-wise CPU updates across all GPU ranks on a node.

What This Paper Is About

Full-parameter fine-tuning of large language models needs roughly 16N bytes of training state for N parameters, which can exceed GPU memory. Host-resident layer streaming solves the capacity problem by keeping persistent states in CPU memory and streaming layers through a small reusable GPU window, but extending this approach from one GPU to many on a shared host creates two new problems: replicated transfers and gradient returns multiply host traffic by the rank count, and strong scaling shrinks the GPU computation window that used to hide host-side work. SlideDP addresses both by maintaining one shared host state, decoupling communication routes from that state's layout, and pipelining parameter delivery, gradient reduction, and CPU updates across ranks and chunks under a GPU memory budget.

Key Contributions

  1. A shared-host runtime for layer-streaming data parallelism. The runtime coordinates parameter delivery, gradient reduction and return, and layer-wise CPU updates inside a bounded GPU working window. It keeps one authoritative host copy of master weights and optimizer states while supporting either replicated or sharded parameter delivery over the same persistent layout. GPU-side reduction returns one aggregated gradient per layer, and chunking overlaps parameter conversion with transfers and gradient return with CPU updates. Per-layer parameter-version and update dependencies preserve synchronous data-parallel semantics.

  2. An analytical step-time model of shared-host scaling. The model expresses a training step in terms of shared-resource service demands and cross-layer, cross-rank execution dependencies, including readiness, buffer reuse, and parameter versions. It characterizes host-bound and GPU-bound execution and explains when reducing traffic shortens a step, when shrinking computation windows expose host work, and why route choices depend on both topology and workload.

  3. Pipeline-aware policy selection with measured cost. AutoPolicy configures communication routes, chunk sizes, and activation layouts for a given workload and hardware. Elastic Checkpointing supplies per-layer choices for retention, offloading, and recomputation. Staged profiling selects routes and chunks, screens activation candidates, and calibrates their costs in a mini-pipeline before assigning policies to layers and validating the configuration.

  4. A working implementation with multi-platform evaluation. SlideDP is implemented in PyTorch Distributed and evaluated on RTX 4090, A800, and H100 platforms, covering 4B–72B models.

Main Findings

  • Throughput over prior systems: In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46–2.64× over SlideFormer, MegaTrain, and ZeRO-Offload.

  • Comparison to GPU-resident training: On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%.

  • Compute-reference efficiency: Across 4B–72B models on four H100s, SlideDP reaches approximately 85–92% of compute-only reference throughput projected from single-GPU measurements.

  • Weak scaling: For Qwen3-32B at 64 sequences per GPU, it maintains 95–98% weak-scaling efficiency across two to eight A800s.

  • Long sequences and large models: On four H100s, SlideDP fine-tunes Qwen3-14B with over 1M tokens per step without gradient accumulation and, separately, with 256K-token sequences. It also fine-tunes Qwen2.5-72B on a PCIe-only workstation with four RTX 4090 GPUs.

  • Host amplification is real and measurable: On an A800 node, the cast-and-copy of an 800M-parameter layer sustains 15.6 GB/s per GPU when one GPU copies, but 8.7 GB/s per GPU when two copy concurrently. On the H100 host, CPU Adam reaches an estimated 182–184 GB/s at 16 physical cores across 101–488M-parameter layers, and increasing to 32 cores improves update speed by only 8.5–10.0%.

  • Strong scaling exposes host work: For Qwen3-14B at a fixed global batch of 256, increasing the A800 GPU count from four to eight reduces the matched compute reference from 27.2 to 14.4 s, while SlideFormer's step time remains near 43 s. SlideDP reduces step time from 29.6 to 15.3 s.

  • The best communication route depends on topology and regime: On a 4× A800 node with NVLink bridges and Qwen3-14B at 16 sequences per GPU, sharded delivery is 17% faster than replicated; after the bridges are removed, replicated is 26% faster. At 8 sequences per GPU the gaps are 33% and 19%. In GPU-bound eight-GPU workloads (64–256 sequences per GPU), the two routes differ by under 2.5%.

  • Spare HBM is a policy variable: At 32 sequences per GPU, the fixed Qwen3-14B configuration uses 15.0 GiB of peak CUDA reserved memory on an 80-GB H100, leaving headroom that can be spent on retaining activations, retaining intermediates, or reducing replay. Because these choices reduce different exposed costs, memory allocation becomes a runtime decision rather than a fixed checkpointing choice.

  • Traffic accounting: The paper tabulates per-layer host traffic for four datapaths (replicated or sharded delivery, each with CPU-side or GPU-side reduction). Replicated delivery with CPU reduction multiplies both H2D and D2H traffic by the rank count R and adds a host-side reduction. Sharded delivery with GPU reduction keeps aggregate parameter and gradient payload independent of R for layers with fixed load counts.

Methodology in Plain English

The researchers start from an existing idea — keep the model's persistent training state in CPU memory, and stream one layer at a time through a small GPU buffer. They then ask what breaks when several GPUs on the same machine all try to do this at once. Their answer is a runtime that treats the host state as a single shared object rather than something each GPU independently replicates.

Concretely, the design separates what state exists from how it travels. The host keeps one copy of the FP32 parameters and the Adam optimizer moments, updated once per iteration by a single worker. Each GPU rank receives either a full copy of a layer's BF16 parameters or a shard of them, reconstructed with an all-gather. Gradients are either summed on the GPUs and returned to the host as one aggregated gradient, or returned per-rank for host-side reduction. Because the route can change without changing state ownership, the system can pick a route that suits the hardware.

Overlapping is handled at two granularities. Completed layers can start their CPU update while earlier layers are still in backward, and within a layer, chunks let the CPU cast one chunk while another chunk is being transferred, and let finished gradient copies feed Adam while later copies are still in flight. Version gates ensure every rank computes a layer with the same parameter version and that a layer's next dispatch waits for its update — keeping the schedule synchronous.

To choose a configuration, the authors build a step-time model from resource service demands (GPU compute, interconnect, CPU update worker, host DRAM, shared PCIe) and dependency constraints (data, ordering, buffer release, parameter version). This model identifies whether a workload is host-bound or GPU-bound. AutoPolicy then measures rather than assumes: it probes candidate delivery and reduction routes, tunes chunk sizes, profiles per-layer checkpointing candidates (full-layer, seven-bit segment masks, selective activation checkpointing, and no-checkpoint variants, each combined with an offload ratio between 0 and 1), runs shortlist candidates through a reduced real pipeline to calibrate their marginal cost, solves an integer program to allocate layer counts under a memory budget, and validates the result with short full-model runs.

Evaluation covers batch-size scaling and CPU memory for Qwen3-8B on four RTX 4090s and Qwen3-14B on four H100s, plus GPU memory versus batch size for Qwen3-14B, alongside route comparisons, GPU-count scaling, and chunk-size sweeps.

Why This Matters

Impact on research. The paper reframes multi-GPU host-resident training as a shared-resource scheduling problem rather than a memory-capacity problem. Its four observations — avoidable amplification, residual pipeline exposure, topology-and-regime-dependent route choice, and the conditional value of spare HBM — give a vocabulary for reasoning about why naive data-parallel extensions of single-GPU streaming underperform. The step-time model, which combines service-demand lower bounds with dependency-graph completion times, is a reusable template for analyzing other offloaded or heterogeneous execution schedules.

Real-world applications:

  • Fine-tuning very large models on workstation-class hardware: Qwen2.5-72B is fine-tuned on a PCIe-only workstation with four RTX 4090 GPUs.
  • Long-context fine-tuning: SlideDP supports 256K-token sequences for Qwen3-14B on four H100s.
  • Large-batch training without gradient accumulation: over 1M tokens per step for Qwen3-14B.
  • Deploying on nodes without fast GPU interconnects: the paper explicitly measures the reversal of the preferred delivery route when NVLink bridges are removed on A800, which matters for PCIe-only or bridge-less clusters.

Industry relevance. Much of the installed base of training hardware is memory-limited relative to the largest models, and multi-GPU nodes share CPU, DRAM, and host–device bandwidth by construction. A runtime that turns spare host memory into an efficient fine-tuning tier, and that adapts its communication policy to the actual machine, addresses a practical constraint for teams that cannot simply add GPUs or rely on GPU-resident sharding.

Future Directions

  • Multi-node scaling. The design and evaluation target a single multi-GPU node with a shared host. How the authoritative host state and cross-rank pipeline extend across nodes, where host resources are no longer shared, is not reported.
  • Interaction with other parallelism strategies. The paper argues data parallelism preserves the layer boundaries that streaming relies on, and notes that tensor parallelism complicates host–device overlap. Whether the routing and scheduling mechanisms compose with tensor or pipeline parallelism is left open.
  • Generality beyond the evaluated configurations. The evaluation covers Qwen-family models from 4B to 72B on RTX 4090, A800, and H100. Behavior for other architectures, optimizer variants beyond Adam, and topologies not measured in the paper is not reported.
  • Cost of policy selection. The paper states that it reports selection overhead and the number of steps needed to amortize it, but the numerical values for those overheads are not present in the available content and would need to be checked in the full paper. Reducing probe cost or making AutoPolicy adaptive during training are natural follow-ups.

Target Audience

Systems researchers and engineers working on distributed training, memory-constrained fine-tuning, or heterogeneous CPU–GPU scheduling will get the most from this paper. It is also relevant to practitioners who need to fine-tune multi-billion-parameter models on limited GPU hardware — including bridge-less multi-GPU nodes and PCIe-only workstations — and to readers interested in measurement-guided systems design, where an analytical model narrows the search space and runtime profiling picks the final configuration.

Authors’ abstract

Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.

Read the original paper