Research
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
Overview Research area: Federated learning for foundation models (Federated Foundation Models), reinforcement learning for LLM reasoning, and privacy-preserving model improvement. Technical level: Adv

- arXiv
- 2602.12014
- Published
- 2026-02-12
- Authors
- Gongxi Zhu, Hanlin Gu, Lixin Fan, Qiang Yang, Yuxing Han
AI summary
Overview
Research area: Federated learning for foundation models (Federated Foundation Models), reinforcement learning for LLM reasoning, and privacy-preserving model improvement.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning for LLMs (PPO, DPO, GRPO), parameter-efficient fine-tuning, and federated learning protocols.
Scope: The paper introduces FedGRPO, a framework that improves a server-side foundation model using only scalar reward signals returned by domain-specialist client devices, avoiding the exchange of model parameters or synthetic data.
What This Paper Is About
Federated Foundation Models aim to combine a server's large general-purpose model with the domain knowledge held on small client devices, but existing approaches send model parameters or synthetic data back to the server, which is expensive in bandwidth and leaks private information. FedGRPO reframes the problem as a reinforcement-learning-style evaluation process: clients act as reward models that score the server's generated answers, and only tiny scalar reward values travel back to the server. The goal is to reach accuracy comparable to centralized training while shrinking communication and privacy exposure.
Key Contributions
-
FedGRPO pipeline. A reinforcement-learning-inspired FedFM pipeline that recasts large-model refinement as a reward-based evaluation process, packaging each query together with its solution rationale into candidate policies and using a lightweight confidence graph built from auxiliary data for competence-based expert selection. This permits parallelized evaluation on resource-constrained clients without local fine-tuning.
-
Federated Group Relative Policy Optimization module. An aggregation scheme that collects only scalar reward signals from selected expert clients and converts them into a group-relative loss, eliminating transmission of raw data, model updates, or high-dimensional representations.
-
Dual evaluation on the client side. A design in which each client uses either answer-based evaluation (exact-answer check against a trusted local ground truth, binary score) or model-based evaluation (a self-trained reward model, real-valued score), selected per question by a gating indicator, so every client uses its strongest available knowledge source.
-
Empirical validation. Experiments across three model scales and two training datasets showing that FedGRPO closely matches or occasionally exceeds centralized GRPO and outperforms model-level and synthetic-data-level federated baselines, while reducing communication by orders of magnitude.
Main Findings
-
Outperforms FedFMs baselines on math. On the MATH-benchmark training set with Qwen2.5-Math-7B, FedGRPO reaches an average accuracy of 0.369 versus 0.253 for the best federated baseline, DPSDA-FL+SFT. With Qwen2.5-Math-1.5B on the same set, FedGRPO scores 0.338 versus 0.275 for DPSDA-FL+GRPO.
-
Approaches or surpasses centralized GRPO. On OpenR1-Math with the 7B model, FedGRPO achieves average accuracy 0.388, exceeding Central-GRPO's 0.379. On MATH-benchmark with the 7B model, FedGRPO reaches 0.369 versus Central-GRPO's 0.370. On OpenR1-Math with Qwen2.5-3B, FedGRPO records 0.227 versus Central-GRPO's 0.225.
-
Works without ground-truth answers. When no client holds ground-truth labels and clients rely purely on model-based evaluation under Non-IID Dirichlet β = 0.1, FedGRPO with Qwen2.5-Math-1.5B achieves an average accuracy of 0.310 versus Central-GRPO's 0.323 and zero-shot's 0.163. For Qwen2.5-Math-7B it reaches 0.327 versus Central-GRPO's 0.369; for Qwen2.5-3B, 0.146 versus 0.168. Fedpetuning and DPSDA-FL are reported as inapplicable in this setting because reference answers are absent.
-
Large communication savings. FedGRPO requires only 2.4 MB of reward-signal transfer for a full training run and stays constant regardless of model size, because it transmits only a short reward value (1.0 or 0.0) per policy. DPSDA-FL requires 102.5 MB (a 40× increase), while FedPETuning scales with model size up to 6.1 GB for the 7B model — a reduction of two to three orders of magnitude.
-
More clients help. As the client count K grows from 4 to 20, average accuracy for Qwen2.5-Math-1.5B rises from approximately 0.29 to 0.36; AMC accuracy rises from about 0.37 to 0.47 and Olympiad accuracy from 0.32 to 0.38. The same trend is observed for Qwen2.5-3B.
-
Gains on multi-hop question answering. On the QA task, FedGRPO with Qwen2.5-3B attains 24.5% average accuracy versus 8.0% for zero-shot and 16.5% for federated SFT. With Qwen2.5-1.5B it attains 20.6% versus 7.0% for zero-shot and 8.4% for federated SFT.
-
Privacy posture, not a privacy proof. The paper's stated claim is that exchanging reward values instead of data or model updates reduces privacy risk under an honest-but-curious threat model, and that it avoids the leakage risks inherent in model-level or synthetic-data-level transfer. No formal differential-privacy guarantee is reported.
Methodology in Plain English
The server holds a large foundation model and a tiny set of auxiliary question-answer pairs (100 samples in the experiments). Each client holds private domain data and its own small trained model.
For any question, the server embeds it with a frozen encoder and retrieves the L most similar auxiliary exemplars (L = 20) by cosine similarity. Those exemplars are broadcast to all clients, and each client reports how accurately it handles them — a "competence score" between 0 and 1. The server keeps only the top-M clients (M = 2) for that specific question, forming a confidence graph of who is good at what.
The server then samples a provisional answer from its own policy, and broadcasts the question-answer pair to the selected experts. Each expert scores the answer using either an exact-answer check against a local ground truth (if it has one) or a locally trained reward model. Only the resulting scalar goes back.
The server standardizes these scores within the group — subtracting the group mean and dividing by the group standard deviation — producing a group-relative reward that is scale-invariant across different evaluation modes and dampens outliers. That reward drives a policy-gradient update on the server model, reinforcing answers that beat the expert-group average. Training runs are configured with a learning rate of 3.0 × 10⁻⁶, temperature 0.7, 8 sampled candidate policies per question, and a maximum generation length of 2048 tokens; 2,000 samples are randomly drawn from each training set, and each result is the average of three repeated experiments.
Why This Matters
Impact on research. The paper offers a third path between model-level and data-level knowledge transfer in federated foundation models, arguing that evaluation scores are orders of magnitude smaller than parameters or synthetic data and far less informative to an adversary. It also extends the "group relative" idea from GRPO into a cross-device federated setting, where the group is composed of clients rather than candidate policies.
Real-world applications:
- Improving a general-purpose medical or legal LLM using feedback from hospitals or firms that cannot share patient or client records.
- Domain-specialized assistants in finance or customer support, where regional branches hold proprietary queries and answers.
- On-device personalization of a shared server model across many heterogeneous phones or edge devices with limited bandwidth.
- Multi-hop question answering over proprietary knowledge bases distributed across organizations.
Industry relevance. The constant 2.4 MB communication cost regardless of model size is directly relevant to deployment economics, and the ability to operate without ground-truth labels broadens the set of clients that can participate. The reduction from 6.1 GB (FedPETuning, 7B) to 2.4 MB is a two-to-three-order-of-magnitude change that matters for real network budgets.
Future Directions
- Formal privacy guarantees. The current exchange of scalar rewards is argued to reduce leakage informally; whether these signals can still be inverted or membership-inferred remains an open question, and adding differential privacy is not explored here.
- Full expert-selection ablation. The paper states that detailed results on varying the number of selected experts M are placed in the Appendix due to space limits; the provided content does not report those numbers.
- Scale and heterogeneity beyond the tested range. Experiments cover 4 to 20 clients, up to 320 communication rounds, and model sizes from 1.5B to 7B; behavior at larger client populations and model scales is not reported.
- Reward-model quality dependence. In the no-answer setting every client trains a Qwen2.5-Math-1.5B reward model, so the method's ceiling depends on how good a weak local evaluator can be; improving reward-model quality or handling adversarial or low-quality experts is not addressed.
Target Audience
Researchers and practitioners working on federated learning for large language models, privacy-preserving machine learning, and reinforcement-learning-based post-training of LLMs. It is most useful to readers who already understand GRPO-style policy optimization and federated aggregation, and who are evaluating alternatives to federated fine-tuning or synthetic-data sharing for bandwidth- and privacy-constrained deployments.
Authors’ abstract
One important direction of Federated Foundation Models (FedFMs) is leveraging data from small client models to enhance the performance of a large server-side foundation model. Existing methods based on model level or representation level knowledge transfer either require expensive local training or incur high communication costs and introduce unavoidable privacy risks. We reformulate this problem as a reinforcement learning style evaluation process and propose FedGRPO, a privacy preserving framework comprising two modules. The first module performs competence-based expert selection by building a lightweight confidence graph from auxiliary data to identify the most suitable clients for each question. The second module leverages the "Group Relative" concept from the Group Relative Policy Optimization (GRPO) framework by packaging each question together with its solution rationale into candidate policies, dispatching these policies to a selected subset of expert clients, and aggregating solely the resulting scalar reward signals via a federated group-relative loss function. By exchanging reward values instead of data or model updates, FedGRPO reduces privacy risk and communication overhead while enabling parallel evaluation across heterogeneous devices. Empirical results on diverse domain tasks demonstrate that FedGRPO achieves superior downstream accuracy and communication efficiency compared to conventional FedFMs baselines.