Research
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
Overview Research area: Natural Language Processing, specifically legal-domain LLM evaluation and benchmark construction for Japanese. Technical level: Intermediate. The methods are straightforward (d

- arXiv
- 2511.22869
- Published
- 2025-11-28
- Authors
- Zhihan Cao, Fumihito Nishino, Hiroaki Yamada, Nguyen Ha Thanh, Yusuke Miyao, Ken Satoh
AI summary
Overview
Research area: Natural Language Processing, specifically legal-domain LLM evaluation and benchmark construction for Japanese.
Technical level: Intermediate. The methods are straightforward (dataset construction from official exam PDFs, binary classification evaluation), but the paper assumes familiarity with LLM benchmarking practice, F1 scoring, and zero-shot/four-shot prompting.
Scope: The paper introduces JBE-QA, a 3,464-item Japanese bar exam question-answering dataset covering the Civil Code, the Penal Code, and the Constitution, and reports baseline results for 26 large language models.
What This Paper Is About
Large language models need accurate legal knowledge to be reliable, but existing Japanese legal evaluation resources focus mostly on the Civil Code and largely ignore the Penal Code and the Constitution. The authors build a new benchmark from the multiple-choice (tantō-shiki) section of the Japanese bar exam from 2015 to 2024, converting each question into individual true/false judgments so that model performance can be measured in a fine-grained way. They then use it to establish baseline scores across a wide range of proprietary, open-weight, Japanese-specialised, and reasoning models.
Key Contributions
- JBE-QA dataset: The first comprehensive Japanese legal-domain benchmark covering the Civil Code, the Penal Code, and the Constitution simultaneously, containing 3,464 question-answer pairs with balanced labels (52.4% False, 47.6% True).
- A decomposed question format: Each original multiple-choice question is split into independent binary classification problems, with structured fields including
id,year,subject,subject_jp,instruction,question,label,answer, and the optionaltheme,lead_in, andremarkfields. - A 26-model baseline: Zero-shot and four-shot evaluations of proprietary, open-weight, Japanese-specialised, and reasoning models, reported with both F1 and a "faithfulness score" (the ratio of instances where a model produces a binary output as instructed).
- A subject-level and qualitative analysis: Per-subject F1 breakdowns plus a case study showing that models rely on superficial phrasing rather than precise knowledge of Penal Code articles and precedent.
Main Findings
- Proprietary models lead: Proprietary models achieve F1 scores ranging from 0.602 to 0.861, except the outliers Haiku-3.5 and Sonnet-4.5. Open-weight models score lower, with F1 ranging from 0.320 to 0.731.
- Reasoning helps, and the best model is Opus 4.1 with reasoning: It reaches 0.814 zero-shot and 0.861 four-shot, the top scores in the paper. GPT-5 follows at 0.794 zero-shot and 0.783 four-shot, and Sonnet 4 with reasoning at 0.780 and 0.776.
- Instruction-following failures can crush scores: Most models have faithfulness scores above 0.990, but Haiku-3.5 has 0.259 and Sonnet-4.5 has 0.027 because they append explanations after the answer, causing non-binary outputs that are scored as 0 by default. Their F1 scores in the zero-shot setting are 0.148 and 0.055 respectively.
- Few-shot examples help proprietary models but hurt open-weight ones: Eight of thirteen open-weight models do not improve in the four-shot setting. Llama-3.3 scores 0.620 four-shot, 0.058 below its zero-shot score. The LLM-jp-3.1 models degrade sharply, with LLM-jp-3.1-13B falling from around 0.582 to around 0.244 and LLM-jp-3.1-8x13B from 0.495 to 0.320.
- The Constitution is the easiest subject: It is always the highest or second-highest scoring subject. Opus-4.1 with reasoning exceeds 0.900 F1 on the Constitution in both settings (0.905 zero-shot, 0.918 four-shot).
- The Civil Code and Penal Code are the hardest: In the zero-shot setting, 11 of 26 models perform worst on the Civil Code and 15 on the Penal Code. In the four-shot setting the split is 14 and 12.
- Reasoning shifts where models struggle: Among seven proprietary non-reasoning models, six perform worst on the Penal Code, whereas among six proprietary reasoning models, four perform worst on the Civil Code.
- Japanese specialisation gives modest gains: Swallow-3.1 and Swallow-3.3 show an advantage of less than 0.03 over their Llama bases; ABEJA-V2 outperforms its base Qwen2.5-32B by 0.065. The largest gain appears for Swallow-3.3 over Llama-3.3 in the four-shot setting, at 0.111.
- Japanese-native pretraining from scratch underperforms: LLM-jp-3.1 models stay below 0.600 zero-shot and below 0.350 four-shot, which is close to or below the 0.488 F1 expected from random guessing on this label distribution.
- One question defeats every model: A Penal Code question on Arson of Uninhabited Buildings (Art. 109, with Art. 115) is unsolved by all models. Opus-4 with reasoning argued that the house was A's own property and inhabited, showing reliance on surface phrasing instead of the relevant article and precedent.
Methodology in Plain English
The authors downloaded publicly available past bar exam PDFs from the Ministry of Justice of Japan covering 2015 to 2024 and converted them into structured XML. They then normalised formats with automated scripts and manually reviewed and corrected extraction errors.
From each original multiple-choice question, they extracted the instructional text, the statement to be judged, and any theme, background text, or clarifying remarks. Rather than asking a model to pick the correct combination of truth values from the original options, they split every question into separate true/false items, so a model judges one statement at a time. Fifty-two questions could not be turned into binary items and were removed: 46 from the Penal Code, 5 from the Constitution, and 1 from the Civil Code.
For evaluation, each model received a Japanese system prompt instructing it to output only 1 (correct) or 0 (incorrect), with no explanation. If a model produced anything other than a binary value, its prediction was recorded as 0. Performance is measured by F1, alongside the faithfulness score. Each model was tested zero-shot and four-shot, with the two positive and two negative exemplars excluded from scoring, so scores are averaged over 3,460 instances. Non-reasoning models ran at temperature 0 with a maximum output length of 1,000 tokens; open reasoning models ran at temperature 1, and proprietary reasoning models used their default temperature, with output length capped at one quarter of the allowed maximum. All evaluations were run once per model.
Why This Matters
Research impact: JBE-QA fills a documented gap in Japanese legal NLP, where prior resources such as COLIEE, the binary questions of Choi et al. (2023), and the Japanese Tort-case Dataset centre on the Civil Code. By covering the Penal Code and the Constitution as well, it enables a more complete picture of what LLMs actually know about Japanese law, and it provides a comparable baseline for future model development.
Real-world applications:
- Bar exam preparation tools that generate and grade practice true/false legal questions for students.
- Legal information systems for non-expert users, where the paper argues that reliable legal knowledge in LLMs reduces expert dependency and improves access to legal information.
- Screening and regression testing of LLMs before they are deployed in legal workflows, using the faithfulness and per-subject scores as diagnostics.
- Comparative model selection for Japanese-language legal products, since the paper quantifies the trade-offs between proprietary, open-weight, and Japanese-specialised options.
Industry relevance: The results give practitioners concrete guidance: reasoning-enabled proprietary models are the strongest option on this benchmark, instruction-following is not guaranteed even for capable models, and adding few-shot examples can backfire for open-weight systems. The per-subject breakdown also shows that strong aggregate performance can hide weak spots in the Civil Code and Penal Code.
Future Directions
- Extending evaluation to ronbun-shiki essay questions, which demand more intensive legal reasoning over longer contexts than tantō-shiki multiple-choice items. The authors note that the difficulty of grading these even for human experts restricts direct implementation as a benchmark.
- Building an automatic update system so that gold labels stay aligned when the underlying laws are revised, since the current dataset is not updated automatically.
- Investigating why few-shot exemplars distract open-weight and Japanese-native models while helping proprietary ones, given the sharp score drops seen for LLM-jp-3.1.
- Improving model handling of questions that require precise knowledge of specific articles and precedent, such as the arson case that no evaluated model solved.
Target Audience
Researchers in legal NLP and Japanese-language LLM evaluation, benchmark builders, and model developers looking for a Japanese legal test set. It is also useful for practitioners selecting or auditing LLMs for Japanese legal applications, and for legal educators interested in automated question generation and assessment. Readers should be comfortable with standard classification metrics and prompt-based evaluation setups.
Authors’ abstract
We introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models' legal knowledge. Derived from the multiple-choice (tanto-shiki) section of the Japanese bar exam (2015-2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions.