Research
Intent Interpretation at RIC Timescales: Jev Decision Models versus Large Language Models in 6G Open RAN
Overview Research area: Intent-based networking and control-loop design in 6G Open RAN (O-RAN), specifically whether the component that translates natural-language intents into A1 policies should be a

- arXiv
- 2609.23136
- Published
- 2026-09-19
- Authors
- Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu
AI summary
Overview
- Research area: Intent-based networking and control-loop design in 6G Open RAN (O-RAN), specifically whether the component that translates natural-language intents into A1 policies should be a typed decision model or a generative large language model (LLM).
- Technical level: Advanced. The paper assumes familiarity with the O-RAN hierarchy (non-RT RIC/rApps over A1, near-RT RIC/xApps over E2), 5G NR scheduling, SLA windows, queueing, and LLM serving economics. The prose is accessible but the technical context is specialized.
- Scope in one sentence: The paper benchmarks three decision models and three hosted LLMs (plus one shared-weights generative reference) as intent interpreters, measures their decision latency against the 1 s near-RT loop budget, and traces the resulting A1/E2 policies through a 21-cell ns-3 simulation and a real srsRAN/Open5GS/OSC near-RT RIC stack to see whether slow interpretation costs control deadlines, RIC capacity, or radio performance.
What This Paper Is About
Intent-based Open RAN needs a component that converts a natural-language intent (for example, "prioritize the emergency-video class in the stadium cluster until the event ends") into a typed A1 policy inside the RIC's control loop. Decision models such as Jev-1.13.0 return typed policy fields directly, while generative LLMs emit the policy token by token, which is slower. The paper asks whether that extra interpretation delay costs control deadlines, RIC capacity, or radio performance — and finds that it costs deadlines and RIC capacity, but that no general radio penalty of interpreter choice was resolved at the base operating point.
Key Contributions
- RANIntent v1 benchmark. A benchmark that types operator and tenant intents to the fields of an A1 policy, with each intent paired to a KPM telemetry table of 3 to 57 cells in fresh, stale, noisy, or contradictory form. Label tuples are fixed before the intent text is written, and a blind verifier checks every intent text against its 21-cell fresh telemetry.
- O-RAN-aligned closed-loop evaluation. A multi-cell NR simulation in ns-3 with 5G-LENA driven by each interpreter's recorded latency and policy, where UEs move and hand over between 21 cells. Three arms — latency-only (L), accuracy-only (A), and net (N) — separate the radio effect of interpretation latency from that of policy errors.
- Real-stack control-path measurement. Live interpretations sent over A1 and E2 to a software gNB on a real stack of srsRAN, Open5GS, and an O-RAN Software Community near-RT RIC, with the path delay from intent issue to the gNB control acknowledgement decomposed into its parts.
- Decision models versus hosted LLMs on one policy schema. Three decision models (Jev-1.13.0, SemIf-Qwen3.5-4B, AnyJev from Nokia Applied Research) compared with three hosted LLMs (DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash), plus the generative reference Qwen3.5-4B-JSON, which shares the weights of SemIf-Qwen3.5-4B and AnyJev.
Main Findings
- The interpreter decides the RIC placement. On the real A1 and E2 path, median interpretation takes 0.286 to 2.35 s, whereas A1 transfer takes 15.5 to 19.1 ms and E2 control 2.4 to 3.1 ms. Hosted Jev-1.13.0 meets the 1 s near-RT budget on 99.8% of calls with a p99 of 0.472 s; GLM-5.3-Flash and Qwen3.8-Flash meet it on 17.9% and 0% and therefore belong in the non-RT loop.
- Decision time dimensions the RIC. At 2 intents/s, Jev-1.13.0 occupies at most 14.9% of its interpretation slots, whereas AnyJev-L0 and Qwen3.8-Flash exceed full utilization and their p95 queue waits reach 34.5 s and 20.3 s.
- Radio conditions outweigh the interpreter. In the 21-cell closed loop, UE speed and intent rate move the affected-class SLA violation by up to 17 percentage points, while ideal enforcement moves it by 3.96 percentage points in the direction each intent requests, against no update at the base point. No hosted LLM showed a resolved increase over Jev-1.13.0 at that point, and a general radio penalty of interpreter choice was not resolved.
- Direct per-cell control gives no resolved SLA gain. Calling the interpreter every second to set per-cell priorities gave no resolved SLA reduction over the policy-and-xApp split, and five of 28 contrasts were resolved increases.
- Accuracy has to be judged at the loop deadline. On 57-cell tables, GLM-5.3-Flash and DeepSeek-V4.1-Flash reach 0.967 and 0.960 against 0.907 for Jev-1.13.0, although neither paired difference is resolved — and GLM-5.3-Flash still meets the near-RT budget on only 17.9% of calls. At that table size Jev-1.13.0 is the least expensive hosted interpreter, at 0.207 USD per 1,000 correct policies.
- Two interpreters saturate their queues at 2 intents/s, and no radio penalty of slow interpreters was resolved at the base point.
Methodology in Plain English
The authors built a benchmark where the correct answer is fixed first — the intended class, scope, priority and other A1 policy fields — and only then is a natural-language intent written to express it. A separate model, in a separate session and without seeing the answer tuple, reads each intent plus its telemetry table and must recover the same labels; only cases that survive that blind check, linting on wording, and length and redundancy rules are kept. Every condition crosses four telemetry sizes (3, 7, 21, and 57 cells) with four telemetry qualities, giving 16 conditions of 300 test cases each, or 4,800 test cases per interpreter. Half the cases depend on reading telemetry to find the target cluster, which the intent text never names; the other half name their cluster directly. Stale tables must be filtered by measurement age, noisy tables carry a bounded noise that cannot flip the answer, and contradictory tables must privilege E2 KPM rows over O1 performance-management rows — and a decoy row appears in 80% of the stale and contradictory tables to catch readers who ignore the rule.
Each interpreter's returned policy is scored twice. First against the typed ground truth, on target-cluster accuracy, full-policy match, schema validity, unsafe rate, latency, tokens and fees. Second in a closed-loop radio simulation: the interpreter's real recorded latency and its actual policy are injected into an ns-3/5G-LENA multi-cell NR network where UEs move and hand over between 21 cells, with an O-RAN-aligned control path carrying A1 and E2 semantics. Three arms isolate causes — L carries only the interpreter's latency with an ideal policy, A carries only the interpreter's policy with an ideal latency, and N carries both — and each run is compared against oracle, fixed-latency, and no-update controls. Outcomes are the affected-class and network-wide SLA violation (primarily over a 2 s window, with 1, 2, 5 and 10 s also reported), cell-edge and mean throughput, handover rate and interruption time, radio link failures, and PRB utilization.
A separate real-stack experiment sends live interpretations over A1 and E2 to a software gNB on srsRAN, Open5GS and an O-RAN Software Community near-RT RIC, breaking the delay from intent issue to control acknowledgement into decision latency, A1 transfer, and E2 control. The paper formalizes the near-RT feasibility of an interpreter as the probability that decision latency plus E2 control delay stays within 1 s, and formalizes "stale-policy exposure" as the user-seconds during which traffic is still served under the previous weights while a policy is being interpreted and delivered.
Why This Matters
The paper reframes a common LLM-for-networking question. Rather than asking whether an LLM can produce a correct policy, it asks whether the correct policy arrives in time for the loop that must apply it, and it separates those two questions experimentally. It also proposes and evaluates a placement the O-RAN specifications do not define: running the interpreter inside the near-RT RIC (mode X), where decision latency counts directly against the 1 s budget.
Real-world applications:
- Stadium and event-time RAN control. The motivating scenario is giving an emergency-video class priority in a stadium cluster for the duration of an event, while UEs move between cells and competing traffic shifts cell load.
- Multi-tenant network exposure APIs. Half the benchmark intents come from four tenants owning subsets of the eight traffic classes, matching how a tenant would request priority changes through an exposure API.
- RIC capacity planning. The queueing results (slot utilization, p95 queue waits, saturation at 2 intents/s) tell an operator how many interpreters, or how much compute, a given intent arrival rate requires.
- Cost-aware interpreter selection. The per-1,000-correct-policies figure gives an operator a way to compare a decision model against a hosted LLM on a basis that includes both accuracy and the loop deadline.
Industry relevance: the result that no general radio penalty of interpreter choice was resolved at the base point, while placement and capacity penalties clearly were, suggests that the choice of interpreter is primarily a control-plane and budget decision rather than a radio-performance decision. That matters to operators deciding whether an LLM-based intent agent can sit in the near-RT loop or must be confined to a non-RT rApp.
Future Directions
- Testing mode X beyond a proposal. The paper notes that O-RAN specifications define no intent ingress at the near-RT RIC, so the xApp-hosted placement evaluated here is a proposal; a standards-side question is whether such an ingress should be defined, and under what budget rules.
- Extending the real stack beyond feasibility. The real-stack evaluation measures the control path with 30 clean intents per interpreter; extending it to live traffic, mobility, and multi-cell scale would test whether the path decomposition holds under load.
- Raising the intent arrival rate. Queue saturation appears at 2 intents/s for two interpreters, and the load-and-scale question (RQ3) is labelled exploratory by the authors. Higher rates, different scheduling disciplines, or batching and caching of interpretations are natural continuations.
- Resolving the accuracy-versus-deadline gap. On 57-cell tables the hosted LLMs reach 0.967 and 0.960 against 0.907 for Jev-1.13.0, but neither paired difference is resolved and the LLMs still miss the near-RT budget on most calls. Whether accuracy can be improved without losing the deadline — for example through tighter readout of a shared language model, as the AnyJev and SemIf-Qwen3.5-4B variants attempt — remains open.
Target Audience
Researchers and engineers working on O-RAN RIC control loops, intent-based networking, and LLM-driven network automation; operators and vendors evaluating where to place an intent interpreter (non-RT rApp versus near-RT xApp); and benchmarking or systems researchers interested in how to score a language model against a hard control deadline rather than against accuracy alone. It is also relevant to anyone studying the cost, queueing, and capacity behavior of hosted LLM inference in operational telecom settings, since the paper prices interpreters in USD per 1,000 correct policies and reports slot utilization alongside radio KPIs.
Authors’ abstract
Intent-based Open RAN needs an interpreter that turns intents into A1 policies within the loop of the RAN intelligent controller (RIC). Decision models such as Jev-1.13.0 return typed policy fields, whereas generative large language models (LLMs) produce the policy token by token. We ask whether the extra delay of LLMs costs control deadlines, RIC capacity, or radio performance. We compare Jev-1.13.0 and two other decision models with LLMs on the RANIntent v1 benchmark, in closed-loop ns-3 simulation and on a real A1 and E2 path. Median interpretation takes 0.286 to 2.35 s, against under 25 ms for A1 and E2 transfer. Jev-1.13.0 meets the 1 s near-real-time budget on 99.8% of calls, while two hosted LLMs meet it on 17.9% and 0%. In the radio network, ideal enforcement moves the affected-class service-level agreement (SLA) violation by 3.96 percentage points in the direction each intent requests, against no update at the base point. No hosted LLM showed a resolved increase over Jev-1.13.0 at that point. At the same point, per-second direct control gave no resolved SLA reduction over a numerical xApp. Slow interpreters miss the 1 s budget, and two interpreters saturate their queues at 2 intents/s, whereas no radio penalty of slow interpreters was resolved at the base point.