Research
From Policy to Logic for Efficient and Interpretable Coverage Assessment
Overview Research area: Natural Language Processing applied to healthcare policy and claims adjudication, specifically neuro-symbolic AI that couples neural retrieval with symbolic rule engines. Techn

- arXiv
- 2601.01266
- Published
- 2026-01-03
- Authors
- Rhitabrat Pokharel, Hamid Reza Hassanzadeh, Ameeta Agrawal
AI summary
Overview
Research area: Natural Language Processing applied to healthcare policy and claims adjudication, specifically neuro-symbolic AI that couples neural retrieval with symbolic rule engines.
Technical level: Intermediate. Readers will benefit from familiarity with retrieval-augmented generation, cross-encoders, and rule-based/expert systems, though the paper explains each component.
Scope in one sentence: The paper proposes a neuro-symbolic pipeline that retrieves the specific coverage-governing passages for a procedure code and converts them into executable symbolic rules, with the goal of making policy review cheaper and more traceable for human reviewers.
What This Paper Is About
Healthcare coverage policies are long, subjective documents, and determining whether a given procedure code is governed by a particular policy provision is slow and inconsistent. Large Language Models can read this text but hallucinate, produce inconsistent reasoning, and become expensive when run repeatedly at scale. The paper's goal is not to let a model decide coverage, but to build a support tool that surfaces the governing policy language, organizes it into explicit facts and rules, and produces auditable rationales that a human reviewer can trace.
Key Contributions
- A framework that supports human experts analyzing complex documents by integrating a neuro-symbolic approach (neural components for language processing, symbolic modules for deterministic rule application).
- A coverage-aware retriever that identifies and extracts governing policy language relevant to specific CPT (Current Procedural Terminology) codes, trained on expert-labeled relevance judgments rather than topical similarity.
- A rule-based symbolic system that substantially lowers the cost of reasoning compared to continuous LLM inference, with attribute generation performed once per CPT code and rule generation once per coverage policy.
- The authors state explicitly that the system does not make coverage determinations; human reviewers retain full adjudication authority, and the tool exists to find supporting language for auditable reasoning.
Main Findings
- Retrieval finetuning improves downstream reasoning: Across seven plan documents, the rule-based method with the finetuned retriever achieved an average accuracy of 0.87 and F1 of 0.93, versus 0.85 accuracy / 0.91 F1 with the zero-shot retriever and 0.82 accuracy / 0.89 F1 for the GPT-4.1 baseline. Finetuning improved accuracy by an average of 2.69% and F1 by 1.72% over the zero-shot baseline.
- The abstract reports headline gains of 44% reduction in inference cost and 4.5% improvement in F1 score.
- Cost scale differs by orders of magnitude: In Table 2, processing 11,000 CPT codes cost $4,840 for GPT-5-mini with retrieved text, $9,680 for GPT-4.1 and for o3 with retrieved text, $38,720 for GPT-4.1 given the entire document, and $22 for the rule-based system (both zero-shot and finetuned retriever variants). Per 1,000 CPT codes the rule-based cost was $2.50.
- LLMs with retrieved text are slightly more accurate but much more expensive: GPT-5-mini (finetuned retriever) reached 0.94 accuracy / 0.96 F1, GPT-4.1 (finetuned retriever) 0.92 / 0.95, and o3 (finetuned retriever) 0.94 / 0.96 — but these costs scale rapidly with dataset size and also require the finetuned retriever, adding a one-time setup cost of $2,680.
- Full-document prompting is both worse and costliest: Supplying the entire plan document to GPT-4.1 yielded 0.82 accuracy and 0.89 F1 at $38,720 for 11,000 codes.
- Training economics: The retriever was trained for 2.5 epochs (~48 hours) on a single node with 8 × NVIDIA H100 GPUs (Azure H100 instances at 6.98/hour), a one-time cost of approximately $2,680. No GPU is needed for rule-based inference.
- Interpretability comes from rule tracing: For procedure S9212, the PyKnow engine flagged the
pregnancy_maternity_servicesrule because its conditionsis_pregnancy=Trueandis_maternity=Truematched the procedure's attributes, letting a reviewer see exactly which factors were applied. - Error analysis identified two failure modes: 73.5% of incorrect cases occurred when the correct attribute was not incorporated into rule creation (often when the attribute list is long and later attributes are overlooked), and the remaining 26.5% occurred when an insufficient set of rules was generated, so no rule was triggered.
- Per-plan variation is substantial: Rule-based (finetuned retriever) accuracy ranged from 0.81 on Plan #1 to 0.92 on Plan #5; GPT-4.1 accuracy ranged from 0.77 (Plans #3 and #6) to 0.88 (Plans #1, #4, #7).
Methodology in Plain English
The task is: given a CPT code, its description, and a policy document, produce a reasoning trace that links the code to the relevant policy language. The pipeline has two phases.
Phase one — retrieving the right text. The authors argue that ordinary semantic search is the wrong tool, because what determines coverage is not what a passage is about. Their examples: an insulin pump CPT is thematically close to diabetes self-management and nutrition passages, but the governing rule is more likely a short paragraph under Durable Medical Equipment or a policy rider; a continuous glucose monitoring CPT clusters with general diabetes advice while the real coverage clause is a concise exclusion elsewhere; debridement and wound care codes sit near diabetic foot care content while the decisive language is in surgical necessity sections with prior-authorization requirements.
To fix this, they trained a cross-encoder to score subsections by whether they explicitly govern a CPT's coverage, limitations, or exclusions. Supervision came from an internal annotation platform and roughly 20 certified coding Subject Matter Experts, with an arbitration process for consistency. Across 172 Certificates of Coverage/Summary Plan Documents, this produced over 1.84 million labeled (CPT, subsection, relevance) pairs; 10% was held out for validation and about 1.61 million pairs were used for training.
The model is a LongformerForMultipleChoice fine-tuned from the allenai/longformer-base-4096 backbone, with a 1,536-token context window that fits entire subsections without truncation. Training used the AdamW optimizer (learning rate 2e-5, weight decay 0.01), bf16 mixed precision, and gradient checkpointing, for 2.5 epochs over ~48 hours. The task is framed as contrastive multiple-choice ranking with cross-entropy loss on the positive passage; the authors note this is functionally equivalent to objectives like InfoNCE. Each CPT query is constructed as <CPT>: <lay description> and choices are raw subsection texts.
A cross-encoder is feasible because the candidate pool per plan is small and well-defined — typically fewer than 60 subsections across the relevant "Covered Services" and "Exclusions & Limitations" sections. At inference, passages scoring above a threshold τ (default 0.25) are kept, capped at five from "Covered Services" and five from "Exclusions & Limitations." If nothing clears τ, the paper recommends escalating the case to a human reviewer; a placeholder row is emitted so downstream stages can track completeness.
Phase two — turning text into rules. This uses PyKnow, a Python library for symbolic reasoning, where users define facts, fields (called "attributes" in this paper), and rules. First, attributes are extracted for each CPT code: each attribute is a yes/no property of the procedure (for example is_implant = True), generated by prompting a model with the CPT code, a short description, and the subsections it was matched to. Attributes are framed as yes/no questions and are created only once per CPT code, then reused across new plan documents; 10 plan documents were used at this step. Second, rules are generated per subsection from the coverage text plus the associated attributes, with the attributes constraining rule construction to avoid syntax errors. Third, the PyKnow engine matches a CPT code and its attributes against the rules and passes the triggered rule's attributes to the human reviewer.
Experimental setup. Evaluation used internal, anonymized data of coverage documents, procedure descriptions, and human-generated determinations. The test set is the same 814 CPT codes across 7 separate coverage documents (5,698 codes total), none of which appeared in retriever training or attribute creation. The baseline is GPT-4.1 with vanilla prompting given the CPT code and the entire plan document; o3 and GPT-5-mini appear in an ablation study. Metrics are accuracy and F1.
Why This Matters
The paper targets a setting where reliability matters more than raw model capability: coverage policy review, where a wrong answer has financial and clinical consequences and a reviewer must be able to justify a decision. Its distinctive move is to treat retrieval as a coverage-governance problem rather than a topical similarity problem, and then to push the reasoning burden out of the LLM and into an auditable symbolic engine that does not need a GPU.
Real-world applications:
- Health insurance claims and prior-authorization review, where a reviewer needs the specific passage governing a code rather than a black-box yes/no.
- Coding audits and compliance checks, since the system emits traceable rationales tied back to policy language, supporting documentation of why a determination was reached.
- Scaling across code sets, including the more than 11,000 CPT codes discussed and other code sets such as HCPCS, where per-code attribute generation happens once rather than per query.
- Regulatory or contractual policy analysis more broadly, where long, subsectioned documents encode if-then conditions that a rule engine can operationalize.
Industry relevance: The cost profile is the headline for operational deployment. Rule-based inference costs $22 for 11,000 CPT codes and requires no GPU, against $4,840 to $9,680 for LLM-based approaches with retrieved text and $38,720 for full-document prompting. The one-time $2,680 retriever training cost is amortized as code volume grows, which the authors argue makes the approach favorable at scale even though LLMs with retrieved text score marginally higher.
Future Directions
- Improving attribute coverage and accuracy, which the error analysis identifies as the most impactful path forward, given that 73.5% of rule failures stem from a missing attribute, often one appearing late in a long input sequence.
- Better handling of long contexts so that relevant attributes are not dropped when attribute lists or documents exceed input limits.
- Deepening rule construction so rule sets fully capture policy nuances and do not leave cases where no rule is triggered (the remaining 26.5% of errors).
- Investigating the performance-versus-cost tradeoff in greater detail, which the authors state they plan to address in future work, since LLM-based methods are slightly more accurate with retrieved text but scale in cost with dataset size.
Target Audience
This paper is most useful to applied NLP and machine learning engineers building retrieval or neuro-symbolic systems for regulated domains; healthcare informatics and claims-processing teams evaluating whether LLM inference is affordable at their code volumes; and researchers working on interpretable or auditable AI, particularly those interested in combining expert-labeled cross-encoder retrieval with symbolic rule engines instead of relying on chain-of-thought prompting alone. Readers focused on purely academic benchmarks will find less here, since the evaluation uses internal, anonymized company data and reports accuracy, F1, and cost rather than public leaderboard comparisons.
Authors’ abstract
Large Language Models (LLMs) have demonstrated strong capabilities in interpreting lengthy, complex legal and policy language. However, their reliability can be undermined by hallucinations and inconsistencies, particularly when analyzing subjective and nuanced documents. These challenges are especially critical in medical coverage policy review, where human experts must be able to rely on accurate information. In this paper, we present an approach designed to support human reviewers by making policy interpretation more efficient and interpretable. We introduce a methodology that pairs a coverage-aware retriever with symbolic rule-based reasoning to surface relevant policy language, organize it into explicit facts and rules, and generate auditable rationales. This hybrid system minimizes the number of LLM inferences required which reduces overall model cost. Notably, our approach achieves a 44% reduction in inference cost alongside a 4.5% improvement in F1 score, demonstrating both efficiency and effectiveness.