Research
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
Overview Research area: Natural Language Processing — evaluation and benchmarking of Large Language Models (LLMs) in the public-safety/policing domain. Technical level: Intermediate. The paper combine

- arXiv
- 2601.03553
- Published
- 2026-01-07
- Authors
- Sangyub Lee, Heedou Kim, Hyeoncheol Kim
AI summary
Overview
- Research area: Natural Language Processing — evaluation and benchmarking of Large Language Models (LLMs) in the public-safety/policing domain.
- Technical level: Intermediate. The paper combines LLM benchmarking with a statistical metric-selection pipeline (multiple regression, Spearman rank correlation, t-tests), but the framework itself is described at a conceptual level.
- One-sentence scope: The paper proposes PAS (Police Action Scenarios), a five-stage framework for evaluating LLMs on realistic police tasks, and applies it to a dataset of 75 question-answer pairs derived from 1,602 official Korean police manuals (8,348 curated entries; over 8,000 documents) to benchmark GPT-4, Gemini, and Claude.
What This Paper Is About
Police officers increasingly use LLMs for tasks such as traffic accident analysis, report generation, phishing detection, and criminal investigation, but no evaluation framework exists that is tailored to police operations. Standard evaluations focus on information accuracy, which can miss responses that are not legally "incorrect" yet could still lead to unlawful arrests or improper evidence collection. The paper's goal is to build a scenario-based evaluation framework with domain-specific metrics, expert-validated reference answers, and then measure how well widely used commercial LLMs actually perform on police-work inquiries.
Key Contributions
- The PAS framework. A five-stage, expandable evaluation framework for policing tasks, formally expressed as E_police = f(S, R, G, M, P): Policing Scenario Definition (S), Reference Responses (R), Response Generation (G), Core Evaluation Metrics (M), and LLM Performance Evaluation for Policing (P). Each stage is designed to be adaptable across different police missions and operational contexts.
- A police-domain corpus and evaluation dataset. The authors curated and refined over 8,000 official Korean police documents, starting from 1,602 official police manuals from the Korean National Police Agency, producing 8,348 curated entries structured as {title, content, question}. From these, 75 question-answer pairs were created, spanning 12 work areas, 5 processes, and 5 manual types (per Table 2 categories).
- A two-stage, expert-validated metric selection method. Starting from 15 candidate metrics (the 12 FLASK core metrics, with factuality split into factuality and groundness, plus numerical sensitivity and logical explanation), the authors used stepwise multiple regression followed by Spearman correlation against police expert judgments to reduce the metric set to five final metrics.
- An empirical benchmark of commercial LLMs. Three commercial LLMs (GPT-4, Gemini, Claude) were evaluated zero-shot at temperature 0.8 on the 75 questions, producing 225 responses judged by three automated LLM evaluations and two expert reviews. The paper states that data and prompt templates are released at https://github.com/Heedou/PASFramework.
Main Findings
- Reference answers score high. Manual-based reference answers achieved an overall score of 3.88 out of 5.0, with factuality 4.15, logical correctness 4.14, and harmlessness 4.11 — validating them as a benchmark standard.
- Six of 15 candidate metrics significantly predicted response quality. The regression model showed a strong fit (R² = 0.856, F = 297.457, n = 300, p < 0.001).
- Five metrics survived both filters and form the final set: Logical Correctness, Completeness, Factuality, Logical Efficiency, and Logical Robustness. A sixth candidate, Logical Explanation, was significant in regression (β = 0.108, SE = 0.036, p = .003) but failed the correlation criterion (ρ = 0.203, p = .055) and was excluded.
- LLM-judge/expert agreement was moderate, not strong. Correlation coefficients for the retained metrics ranged from ρ = 0.315 to ρ = 0.560 (all p < .05), with individual values: Logical Correctness ρ = 0.560, Completeness ρ = 0.409, Factuality ρ = 0.330, Logical Efficiency ρ = 0.330, Logical Robustness ρ = 0.315.
- LLMs underperformed significantly, and worse on police-specific metrics. The average skill difference across key metrics was −1.115, compared with −0.828 across all evaluation skills (p < .05).
- Model-wise key-metric gaps: Gemini −0.893, Claude −1.169, GPT −1.283 — each larger than the corresponding all-metric averages of −0.611, −0.880, and −0.994.
- Overall performance gap. The average difference in overall scores was −1.015. Except for readability, every skill scored significantly lower than reference answers (p < .05). The only model-metric combination not significantly below the reference was Gemini's logical explanation.
- Table 5 overall scores: Reference 3.88, GPT-4 2.69, Gemini 3.06, Claude 2.83. Skill scores: Reference 3.72, GPT-4 2.73, Gemini 3.11, Claude 2.84.
- LLMs are strongest where policing demands least. Models performed best on harmlessness (3.91), readability (3.67), and insightfulness (3.09), while the reference answers led on factuality (4.15), logical correctness (4.14), and harmlessness (4.11) — indicating LLM strength in text coherence over factual and procedural precision.
- Lowest individual metric scores. GPT-4 scored lowest on numerical sensitivity (2.12); Claude scored 1.89 on 112 Report; Gemini 2.36 on conciseness. Reference answers were also lowest on metacognition (2.85) and conciseness (2.76).
- Domain-level weaknesses. All models scored low in 112 emergency calls (2.37), traffic (2.55), and security (2.61). Reference answers scored highest in Crime Prevention (4.11) and Violence & Major Crimes (4.05). GPT-4's best domain was crime prevention (3.22); Gemini's and Claude's highest scores were in police organization and operations.
- A qualitative failure mode was identified. In a Figure 3 case study, LLM responses contained no factual errors and could score highly under standard evaluations, yet they omitted manual-based procedures, practical knowledge, and warnings about legal risks such as unlawful arrest.
- LLM-as-a-judge is not a substitute for experts. The paper reports that pairwise preference agreement between LLM judges and Subject Matter Experts has been as low as 60–64% in specialized healthcare fields, and argues that policing requires expert-in-the-loop validation.
Methodology in Plain English
The researchers built their framework around a specific training habit of police officers: mentally rehearsing responses to incidents before a shift begins. They call this scenario "Police Readiness through Operational Reasoning," where the situation is the preparatory state before duty and the action is mentally simulating responses to anticipated incidents.
To create reference answers, they took 1,602 official Korean police manuals covering investigation, law enforcement, traffic control, and emergency response. Most were PDFs with three to four levels of hierarchy, so the authors segmented them by section headers and reformatted them into a standard {title, content, question} structure. With expert guidance, they filtered for quality and produced 8,348 curated entries, then built 75 question-answer pairs by identifying a key question and locating its answer within the same entry.
They then ran zero-shot experiments with GPT-4, Gemini, and Claude at temperature 0.8. Each model received only a brief scenario describing an officer's preparatory state and was told to answer from the perspective of a police officer, without external knowledge or extra context. This produced 225 responses (75 questions × 3 models).
For metrics, they began with the 12 FLASK core metrics, split factuality into factuality and groundness, and added numerical sensitivity and logical explanation, giving 15 candidates. Responses and reference answers were scored on all 15 metrics plus an overall quality score on a 5-point Likert scale, with automated LLM judges at temperature 0.8 plus two expert reviews. The judge prompts included the expert reference response alongside the target response and required reasoning before scoring.
Finally, they filtered metrics in two stages: a stepwise regression kept only metrics significantly predicting overall quality (p < .05), and a Spearman rank correlation kept only those also positively correlated with expert judgments (p < .05). Statistical testing of model performance used t-tests against the reference answers.
Why This Matters
Research impact. The paper argues that existing evaluations of LLMs for police work focus on information accuracy and legal matching or crime classification, which overlooks situational fit, legal-procedural alignment, and judgment rationality. PAS offers a replicable template with expert-validated metrics and a documented dataset, and the authors state the framework is designed to be transferable to other professional fields. It also adds evidence that LLM-as-a-judge scores align only moderately with domain experts in high-stakes settings.
Real-world applications:
- Pre-shift readiness tools that help officers rehearse protocols and recall procedures for anticipated incidents.
- Quality gates for LLM-generated police reports, dispatch suggestions, or investigation guidance before field use.
- Jurisdiction-specific calibration, since the Reference Responses stage lets local experts replace Korean Criminal Procedure Act-based answers with answers tailored to another jurisdiction's procedures.
- Risk screening for legal-procedure violations, such as unlawful arrest or improper evidence collection, that accuracy-only evaluations would not catch.
Industry relevance. The findings indicate that general-purpose commercial LLMs — GPT-4, Gemini, and Claude — do not currently meet police-domain requirements, particularly fact-based recommendations, factual accuracy, and numerical sensitivity. This implies demand for domain-specialized models and for expert-in-the-loop evaluation pipelines rather than fully automated benchmark scores. The authors note that since LLM-based dispatch suggestions have shown regional and racial bias, and police work demands strict legal compliance, premature adoption carries liability risk.
Future Directions
- Expand PAS with real-time data, moving beyond static manual-derived questions toward live operational inputs.
- Broaden jurisdictional diversity. The current evaluation used Korean manuals, which the authors identify as a limitation on general applicability; the framework's Reference Responses stage is the proposed mechanism for adapting to other legal systems.
- Refine evaluation strategies. The moderate correlations (ρ = 0.315 to 0.560) mean LLM judges cannot replace human judgment, so improved judging protocols and stronger expert-in-the-loop designs remain open problems.
- Develop police-specific LLMs. The paper's stated ultimate goal is to determine the direction for building future models aligned with real-world police standards, and it leaves the selection and training of such models to later work.
Target Audience
This paper is most useful to researchers working on domain-specific LLM evaluation, especially those benchmarking models in high-stakes professional fields; to police administrators, policy makers, and police science institutes assessing whether to adopt LLM tools; to developers building legal or public-safety AI systems who need domain-specific evaluation metrics; and to methodologists interested in the two-stage statistical procedure for validating automated metrics against expert judgment. Readers should be comfortable with regression output, correlation coefficients, and t-test reporting.
Authors’ abstract
The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this, we propose PAS (Police Action Scenarios), a systematic framework covering the entire evaluation process. Applying this framework, we constructed a novel QA dataset from over 8,000 official documents and established key metrics validated through statistical analysis with police expert judgements. Experimental results show that commercial LLMs struggle with our new police-related tasks, particularly in providing fact-based recommendations. This study highlights the necessity of an expandable evaluation framework to ensure reliable AI-driven police operations. We release our data and prompt template.