Research
AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation
AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation Overview Research area: Multi-agent systems and large language models applied to semantic web policy generation—specifica

- arXiv
- 2512.00602
- Published
- 2025-11-29
- Authors
- Wanle Zhong, Keman Huang, Xiaoyong Du
AI summary
AgentODRL: A Large Language Model-based Multi-agent System for ODRL GenerationOverview
- Research area: Multi-agent systems and large language models applied to semantic web policy generation—specifically converting natural language data-usage rules into W3C Open Digital Rights Language (ODRL).
- Technical level: Intermediate. Readers need basic familiarity with LLM agent architectures and some awareness of knowledge-graph/semantic web standards (RDF, SHACL, ODRL), but the paper is largely readable without deep background in any one of them.
- Scope in one sentence: The paper introduces AgentODRL, an Orchestrator-Workers multi-agent system with syntax- and semantics-oriented enhancement strategies, evaluates it against prior ODRL generation approaches on a newly built 770-case benchmark, and measures both grammar compliance and semantic fidelity.
What This Paper Is About
Authoring ODRL policies is hard because it requires fluency in RDF graph models, serialization formats, and ODRL's conceptual system—barriers that keep domain experts without a technical background out of the loop. At the same time, high-quality "natural language-to-ODRL" training data is scarce, and existing methods are monolithic: a single LLM must simultaneously parse legal dependencies, segment semantics, and generate output under strict syntax. The goal of this paper is to automate that translation reliably, and to handle policy texts containing complex parallel or recursive logical structures that prior approaches largely avoid.
Key Contributions
- A multi-agent framework for ODRL generation. AgentODRL combines an Orchestrator-Workers architecture with dedicated syntactic and semantic enhancement strategies. The authors report it improves average grammar and semantic scores across all tested models by 5.39% and 14.52% respectively compared to the prior SOTA strategy (SCR-Enhanced), while maintaining near-perfect grammatical compliance.
- A validator-based syntax strategy. A "generate-validate-correct" iterative loop uses a PYSHACL-based validator running SHACL checks against the ODRL information model and vocabulary, feeding error reports back to the Generator agent until the policy passes or a maximum attempt count (set to 8) is reached.
- A LoRA-based semantic reflection mechanism. A lightweight LLM (Qwen3-4B-Instruct(BF16)) is fine-tuned with LoRA (r=16, alpha=32, 2,380 synthetic samples, 3 epochs, single NVIDIA 4090 GPU) into an expert extractor that produces a "semantic checkpoint" list, against which the Generator must validate its output.
- A new benchmark dataset. The paper constructs and releases 770 use cases, comprising 440 simple cases, 220 parallel-relationship cases, and 110 sequential-dependency cases, all situated in data space contexts, with each case pairing a natural language policy against its ODRL target.
Main Findings
- The full pipeline outperforms prior strategies. AgentODRL's Full Pipeline (AOFP) improves average grammar scores by 5.39% and average semantic scores by 14.52% across the three GPT-4.1 series models relative to SCR-Enhanced. Peak gains reached 7.32% on grammar (GPT-4.1) and 28.67% on semantics (GPT-4.1-nano).
- Grammar scores approach perfection under the validator loop. Within AOFP, grammar scores for all models rose to near-perfect levels (mostly exceeding 99), which the authors attribute to the closed-loop correction mechanism rectifying syntactic hallucinations.
- Smaller models benefit most. With the baseline strategy, GPT-4.1-nano scored 34.87 on semantics for recursive structures; AOFP raised this to 61.53, described as a 76.46% improvement. Model capability still sets the performance ceiling (GPT-4.1 > GPT-4.1-mini > GPT-4.1-nano).
- Performance is stable as complexity increases. Going from simple to recursive cases, the Ontology-Guided Strategy degrades sharply, while AOFP stayed high—for example, GPT-4.1's semantic score only dipped from 98.64 to 96.40.
- Reflection counts scale with difficulty. Weaker models and complex cases required more iterations: GPT-4.1-nano averaged 7.32 reflections on recursive cases, versus 1.22 for GPT-4.1 on simple cases.
- Workflow path should match task complexity. On GPT-4.1-nano, adding the Splitter before the Generator raised semantic scores from 69.44 to 84.88 on parallel structures; the Rewriter→Splitter→Generator path achieved the highest semantic score of 82.00 on recursive cases and 88.07 across all use cases.
- The Orchestrator trades a little accuracy for efficiency. Its average semantic score across all use cases was 80.22, above the plain ODRL Generator's 72.35 but below the best fixed paths (84.02 for Splitter-based, 88.07 for Rewriter-based). It consumed 46,209,570 tokens, fewer than the Splitter path (47,917,323) or the Rewriter path (49,451,650), and more than the plain Generator (33,856,708).
Methodology in Plain English
Rather than asking one model to do everything, the authors split the job the way a team would split it.
First, a central Orchestrator agent reads the incoming natural language policy and classifies its structural complexity using a taxonomy the authors define: simple cases (one self-contained policy), parallel cases (multiple relatively independent policies in one text), and recursive cases (cross-clause dependencies, such as a penalty in one article that hinges on a violation of another). Each rule is formally described as a tuple of party, asset, action, policy clauses, and contextual constraints.
Second, the Orchestrator dispatches the work to specialized Worker agents. A Rewriter resolves cross-references by inlining the referenced clause into the referring clause—a "structure-preserving inlining" that removes the dependency while keeping clauses structurally separate. A Splitter breaks parallel use cases into independent policy units, initiating a new unit only when the core asset, the assigner-assignee relationship, or the policy's fundamental purpose changes, and then assigning an ODRL type (Agreement, Offer, or Set) by heuristics. A Generator converts each well-structured unit into ODRL.
Third, the Generator's output quality is policed in two ways. A PYSHACL-based validator checks each generated policy against SHACL constraints derived from the ODRL standard and returns an error report if it fails, prompting the model to revise—repeating until it passes or hits the attempt limit. Separately, the LoRA-finetuned small model extracts the key semantic elements into a checklist, and the Generator must confirm its policy encodes every point.
Evaluation uses two custom metrics. The Grammar Score is computed as (1 − errors/constraints) × 100% from SHACL validation. The Semantic Score extracts a ground-truth checkpoint list from the original text using an "Identifier" LLM, then has a Jury of two independent LLMs (K=2) score how faithfully each semantic unit is reflected. All configurations were run three times and averaged. The dataset itself was bootstrapped from 70 hand-written seed cases (40 simple, 20 parallel, 10 sequential-dependency) sourced from academic literature, an industry standard draft, Creative Commons 4.0, GDPR, and CCPA, then augmented with Gemini 2.5 Pro under principles of core logic preservation and contextual element transformation, followed by manual review.
Why This Matters
Impact on research. The paper argues that ODRL generation is not one task but several cognitively distinct sub-tasks, and that architectural decomposition—not just a bigger model—is a viable answer. It also fills a concrete gap: the authors state that no ODRL generation dataset previously existed, so the 770-case benchmark is itself a research contribution, and it deliberately includes the parallel and recursive structures that prior work avoided.
Real-world applications.
- Trusted data spaces: IDSA-style ecosystems where participants must express machine-readable usage policies for shared data assets.
- Regulatory compliance modeling: translating provisions from instruments such as the GDPR, the EU AI Act, and the CCPA into enforceable policies.
- Standard and open licensing: converting website terms of service, standards documents, and licenses such as Creative Commons 4.0 into ODRL.
- Commercial agreements: letting business and legal staff describe data-sharing restrictions in ordinary language rather than in RDF.
Industry relevance. ODRL is the W3C standard adopted by initiatives such as the International Data Spaces Association for describing usage policies over data assets. The paper's framing is that the technical barrier of Semantic Web tooling blocks non-technical domain experts from participating, and that lowering this barrier matters for cross-entity data sharing under data sovereignty and interoperability principles. The Orchestrator's cost/accuracy trade-off is also directly relevant to deployment economics: it used fewer tokens than the higher-scoring fixed workflows.
Future Directions
- Close the routing gap. The Orchestrator averaged 80.22 in semantics versus 88.07 for the best fixed path, so better complexity classification and routing could recover that headroom without giving up the token savings.
- Extend the complexity taxonomy and dataset. The benchmark currently covers three structural forms in data space scenarios; whether the approach transfers to other policy domains, languages, or longer regulatory texts is not reported.
- Push semantic checking beyond a checklist. The current semantic strategy relies on extraction of key elements into a checkpoint list; whether richer or more formal semantic validation would help is an open question.
- Reduce refinement cost. The reflection counts (up to 7.32 on average for weaker models on recursive cases) and the 46.2M-token consumption suggest room to make the correction loops cheaper.
- Validate scoring rigor. The Semantic Score depends on an LLM Jury with K=2; the paper does not report human agreement studies against those jury scores.
Target Audience
Researchers and practitioners working on LLM-based multi-agent systems, semantic web and knowledge graph construction, and policy/rights automation will get the most from this paper. It is also relevant to data space architects, standards bodies, and compliance or legal-engineering teams who need machine-readable data usage policies, and to anyone interested in how far task decomposition and self-correction can substitute for larger models on tightly constrained generation tasks.
Authors’ abstract
The Open Digital Rights Language (ODRL) is a pivotal standard for automating data rights management. However, the inherent logical complexity of authorization policies, combined with the scarcity of high-quality "Natural Language-to-ODRL" training datasets, impedes the ability of current methods to efficiently and accurately translate complex rules from natural language into the ODRL format. To address this challenge, this research leverages the potent comprehension and generation capabilities of Large Language Models (LLMs) to achieve both automation and high fidelity in this translation process. We introduce AgentODRL, a multi-agent system based on an Orchestrator-Workers architecture. The architecture consists of specialized Workers, including a Generator for ODRL policy creation, a Decomposer for breaking down complex use cases, and a Rewriter for simplifying nested logical relationships. The Orchestrator agent dynamically coordinates these Workers, assembling an optimal pathway based on the complexity of the input use case. Specifically, we enhance the ODRL Generator by incorporating a validator-based syntax strategy and a semantic reflection mechanism powered by a LoRA-finetuned model, significantly elevating the quality of the generated policies. Extensive experiments were conducted on a newly constructed dataset comprising 770 use cases of varying complexity, all situated within the context of data spaces. The results, evaluated using ODRL syntax and semantic scores, demonstrate that our proposed Orchestrator-Workers system, enhanced with these strategies, achieves superior performance on the ODRL generation task.