Research
Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
Overview Research area: Machine learning / AI safety — specifically content-moderation "guardrail" models that sit alongside large language models to filter harmful inputs and outputs. Technical level

- arXiv
- 2512.05339
- Published
- 2025-12-05
- Authors
- Mahesh Kumar Nandwana, Youngwan Lim, Joseph Liu, Alex Yang, Varun Notibala, Nishchaie Khanna
AI summary
Overview
- Research area: Machine learning / AI safety — specifically content-moderation "guardrail" models that sit alongside large language models to filter harmful inputs and outputs.
- Technical level: Intermediate. The paper assumes familiarity with instruction fine-tuning, LoRA, chain-of-thought data, and F1-based benchmark comparison, but explains its pipeline in accessible terms.
- Scope: The paper describes Roblox Guard 1.0, an 8B-parameter guardrail model instruction-tuned on 384,233 examples to generalize to unseen safety taxonomies, plus a new 2,872-example benchmark called RobloxGuard-Eval.
What This Paper Is About
Existing guardrail models are trained against a fixed, pre-defined list of safety categories, which makes them brittle when a platform's definition of "harmful" changes or differs from the one they were trained on. The authors argue that safety definitions are context-dependent — content like dating talk may be fine for an 18+ product but not a youth platform — so a guardrail should be able to infer and apply a new policy at inference time rather than require retraining. Their goal is a single fine-tuned model that adapts to previously unseen taxonomies and stays robust on out-of-domain safety benchmarks.
Key Contributions
- Roblox Guard 1.0 — an instruction fine-tuned guardrail built on the Llama-3.1-8B-Instruct backbone, released at github.com/Roblox/RobloxGuard-1.0, which the authors report achieves state-of-the-art or comparable performance on several safety benchmarks while generalizing to unseen taxonomies.
- A large, transparent training corpus — 384,233 total examples (213,289 positive and 170,944 negative) drawn entirely from open-source and synthetic sources, augmented with chain-of-thought (CoT) rationales and an "input inversion" technique.
- RobloxGuard-Eval — a public evaluation benchmark (huggingface.co/datasets/Roblox/RobloxGuard-Eval) of 2,872 examples with responses hand-labeled by policy experts, built to stress-test models on fine-grained, platform-specific safety categories.
- A three-stage synthetic data generation pipeline that produces adversarial prompts and responses and labels them with LLM judges calibrated against GPT-4o.
Main Findings
- Strong prompt-level results: Roblox Guard 1.0 scores 91.9% F1 on Aegis 1.0 Prompt (vs. 74.8% for LlamaGuard3, 89.4% WildGuard-7B, 88.7% ShieldGemma-7B, 90.4% BingoGuard-8B, 83.2% GPT-4o, 89.8% NemoGuard-8B) and 89.5% on WildGuard Prompt, the highest among the compared models on both.
- Response-level competitiveness: It reaches 87.3% F1 on BeaverTails (vs. 69.7% LlamaGuard3, 84.4% WildGuard-7B, 84.8% ShieldGemma-7B, 86.4% BingoGuard-8B, 83.8% GPT-4o, 77.6% NemoGuard-8B) and 86.0% on Aegis 2.0 Response — though NemoGuard-8B scores slightly higher there at 87.6%.
- Not uniformly best: On OAI Mod the model scores 70.3%, below ShieldGemma-7B (82.1%), LlamaGuard3 (79.4%), BingoGuard-8B (77.9%), and NemoGuard-8B (77.0%). On XSTest it scores 86.4%, below WildGuard-7B (94.4%), BingoGuard-8B (94.9%), ShieldGemma-7B (92.5%), GPT-4o (90.2%), and LlamaGuard3 (88.3%).
- Large out-of-domain margin on RobloxGuard-Eval: Roblox Guard 1.0 scores 79.6% F1, versus 3.5% for LlamaGuard3-8B, 15.1% for WildGuard-7B, 55.5% for ShieldGemma-7B, 25.9% for BingoGuard-8B, 66.3% for GPT-4o, and 23.6% for NemoGuard-8B — the paper notes most existing models "struggle significantly" and some fall below 30% F1.
- Out-of-domain generalization elsewhere: 79.1% on Toxic Chat, 69.9% on SafeRLHF, 86.4% on XSTest, and 85.7% on HarmBench — datasets with novel prompt styles and harm categories not seen in training.
- Synthetic data is the largest single driver: Removing it collapses RobloxGuard-Eval performance from 79.6% to 20.3%, drops OAI Mod from 70.3% to 49.2%, and WildGuard Response from 80.6% to 68.3%. The no-synthetic variant also failed to produce valid inferences on HarmBench and therefore has no score there.
- CoT matters for reasoning-heavy tasks: Removing CoT lowers Aegis 2.0 Response by 4.4 percentage points and Harmbench by 3.9 points, while slightly improving SafeRLHF (69.9% to 71.0%) and RobloxGuard-Eval (79.6% to 82.3%). The authors hypothesize these sets contain more straightforward violations where reasoning helps less.
- Input inversion helps robustness: Removing it causes a 3.0-point drop on XSTest and a 2.6-point drop on WildGuard Response, which the authors attribute to those benchmarks testing brittleness, over-refusals, and adversarial formats. Notably, the no-input-inversion variant scores higher than the full model on Aegis 2.0 Prompt (88.0% vs. 87.9%), OAI Mod (71.3% vs. 70.3%), and RobloxGuard-Eval (80.7% vs. 79.6%).
- Deployment latency: Served with vLLM on an AWS g6.12xlarge, the model averaged 869.9 ms over 10 runs for a 790-token payload (770 prompt tokens, 20 completion tokens).
- Benchmark saturation: The paper argues existing LLM safety benchmarks are "more likely than not already saturated," motivating the release of a harder, taxonomy-rich set.
Methodology in Plain English
The team started from an off-the-shelf instruction-tuned model, Llama-3.1-8B-Instruct, and adapted it with LoRA (rank r = 16, via the PEFT library) rather than full retraining. Training ran for 3 epochs at a learning rate of 1×10⁻⁴, batch size 8 per device, warmup ratio 0.03, and context length 2408 tokens, in bfloat16 mixed precision on a single machine with 8× A100 GPUs (80GB each). The resulting model is named Llama-3.1-8B-Instruct-RobloxGuard-1.0.
Training data came from two places. Public sources — Aegis, WildGuard, and BeaverTails — contributed examples with their original labels retained, so label granularity from each source survives. The rest was synthetic, made by a three-stage pipeline: first, DeepSeek-R1-Distill-Qwen-7B is handed a policy document covering all safety categories and asked to invent adversarial attack scenarios (system prompts and user messages); second, those scenarios are sent to a mix of models — Mistral-7B-v0.1, Llama-3.2-3B-Instruct, and Qwen2.5-7B-Instruct-Abliterated-v2 — to produce candidate responses; third, LLM judges (Mistral-Small-24B-Instruct-2501 and DeepSeek-R1) read the full exchange and output a binary violation label. These judges were calibrated on a holdout set against human expert labels using GPT-4o as a reference point, where GPT-4o itself scored 85.61% F1, 9.34% false positive rate, and 90.36% recall.
Two design choices push the model toward taxonomy adaptation rather than memorization. Chain-of-thought rationales, generated by DeepSeek-R1 using each dataset's own taxonomy and definitions as the prompt, are included alongside labels. "Input inversion" then permutes the ordering of the target output components — chain-of-thought, label, and category — so the model cannot overfit to a single output format. The instruction set follows a FLAN-style multi-task approach, treating each taxonomy category as its own task. The authors note they deliberately did not unify all datasets under one shared taxonomy.
For evaluation, the team curated RobloxGuard-Eval with internal red-teaming: prompt-response pairs were labeled by 3 policy experts, with agreement from 2 of 3 required for inclusion. Table 1 lists 25 content safety categories, while Table 2 reports example counts for 23 categories (1,980 of the 2,872 examples are "None"); the paper text refers to both 23 and 25 categories in different places.
Why This Matters
This work argues that the standard approach to guardrails — one model, one frozen list of harms — is a structural limitation, not a tuning problem. By showing a single 8B model can generalize to taxonomies it never trained on, the paper provides evidence that context-aware moderation is achievable without per-customer retraining. The ablation showing a 79.6% → 20.3% collapse when synthetic data is removed is a concrete argument that open datasets alone under-cover real-world policy nuance.
Real-world applications:
- Youth-oriented platforms and apps that need stricter definitions of acceptable content than a general-purpose moderation API provides.
- Products with age-tiered or region-specific policies, where the same category (for example, romantic content or dating) must be handled differently by audience.
- Platforms with proprietary, evolving policy suites covering areas like deceptive monetization, off-platform solicitation, paid random items, and advertising practices that general benchmarks rarely test.
- Real-time moderation pipelines in interactive or user-generated-content environments, where the reported 869.9 ms latency for a 790-token payload may be relevant to feasibility.
Industry relevance: the paper targets a practical gap — companies increasingly have bespoke safety taxonomies but no efficient way to get a guardrail model aligned to them. Releasing both the model weights and the evaluation set (the latter under CC BY 4.0 on Hugging Face) gives teams a reproducible baseline, and the reported comparative numbers against LlamaGuard3, WildGuard, ShieldGemma, BingoGuard, NemoGuard, and GPT-4o give a concrete reference frame for build-vs-buy decisions.
Future Directions
- Resolving and consolidating the taxonomy scope of RobloxGuard-Eval, since the paper lists 25 categories in Table 1 but reports per-category counts for 23 in Table 2, and describes the benchmark as covering both 23 and 25 categories.
- Reducing dependence on LLM judges. The synthetic corpus rests on judge models calibrated against GPT-4o, whose agreement with human labels was measured at 85.61% F1 and a 9.34% false positive rate; the paper does not report how judge error propagates into final model behavior.
- Broadening the adaptability claim beyond English text. The motivation cites regional regulations and cultural norms as reasons taxonomies must adapt, but the paper reports no multilingual or multi-region evaluation.
- Extending to other modalities. HarmBench is described as including a multimodal functional category, yet Roblox Guard 1.0 is evaluated only on text.
- Reconciling ablation trade-offs. Input inversion and CoT removal each improved results on some benchmarks and hurt on others (for example, no-input-inversion scored higher on OAI Mod and RobloxGuard-Eval); the paper leaves open how to tune these techniques for a given target domain.
Target Audience
AI safety and trust-and-safety engineers building production moderation systems, researchers working on LLM guardrails and content classification, and platform policy teams who need models aligned to custom taxonomies. It is also useful for evaluation researchers, since the paper makes a direct argument about benchmark saturation and offers a new dataset, and for practitioners weighing an 8B open model against commercial alternatives such as GPT-4o. Readers without a background in fine-tuning will still follow the high-level argument, but the benchmarking and ablation sections assume comfort with F1 scores and benchmark methodology.
Authors’ abstract
Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose risks to users. This challenge underscores the need for robust safeguards that operate across both model inputs and outputs. In this work, we introduce Roblox Guard 1.0, a state-of-the-art instruction fine-tuned LLM designed to enhance the safety of LLM systems through comprehensive input-output moderation, using a pipeline of LLMs to enhance moderation capability. Built on the Llama-3.1-8B-Instruct backbone, our model is instruction fine-tuned to generalize across previously unseen safety taxonomies and demonstrates strong performance on out-of-domain safety benchmarks. The instruction fine-tuning process uses a mix of synthetic and open-source safety datasets, augmented with chain-of-thought (CoT) rationales and input inversion to enhance contextual understanding and decision making. To support systematic evaluation, we also release RobloxGuard-Eval, a new benchmark featuring an extensible safety taxonomy to assess the effectiveness of LLM guardrails and moderation frameworks.