Skip to content
AI.info

Research

FISCAL: Financial Synthetic Claim-document Augmented Learning for Efficient Fact-Checking

FISCAL: Financial Synthetic Claim-document Augmented Learning for Efficient Fact-Checking Overview Research area: Natural language processing, fact verification, and applied generative AI for finance

FISCAL: Financial Synthetic Claim-document Augmented Learning for Efficient Fact-Checking
arXiv
2511.19671
Published
2025-11-24
Authors
Rishab Sharma, Iman Saberi, Elham Alipour, Jie JW Wu, Fatemeh Fard

AI summary

FISCAL: Financial Synthetic Claim-document Augmented Learning for Efficient Fact-Checking

Overview

  • Research area: Natural language processing, fact verification, and applied generative AI for finance (financial claim verification, synthetic data generation, parameter-efficient fine-tuning).
  • Technical level: Intermediate. The paper is readable without deep mathematical background, but familiarity with fact-checking (entailment classification), LoRA fine-tuning, and standard classification metrics (precision, recall, F1, accuracy) helps.
  • Scope: The paper builds a modular pipeline that synthesizes financial claim–document–label triplets (FISCAL-data) and uses them to fine-tune a 7B parameter verifier (MiniCheck-FISCAL) that competes with models many times its size.

What This Paper Is About

Financial applications of large language models need to be factually reliable, but models often hallucinate numbers, dates, and entities, while the most accurate systems are expensive and slow to run at scale. This paper asks whether a small, efficient, open model can be made trustworthy for checking numerical financial claims if it is trained on the right kind of synthetic, domain-specific data. The authors build a data generator (FISCAL) and use its output to fine-tune a lightweight 7B verifier (MiniCheck-FISCAL), then test whether it can match or beat far larger models on financial fact-checking benchmarks.

Key Contributions

  1. FISCAL, a modular synthetic data generator: A pipeline that takes real financial documents and produces labeled claim–document–label triplets by extracting numerical claims and then applying six different perturbation modules (Claim Paraphraser, Conflict Insertion, Fact Exclusion, Fact Value Distortion, Mis-attribution, and Summarization) to create challenging positive and negative examples.
  2. FISCAL-data, a released dataset: 14,304 training triplets, 1,792 evaluation samples, and 1,784 test benchmark samples built from unseen documents in FinanceBench, with no duplicate samples across splits to prevent leakage.
  3. MiniCheck-FISCAL, a compact verifier: A 7B model fine-tuned from MiniCheck-7B using LoRA, with claim verification reformulated as a causal language modeling task that emits a single "yes" or "no" token, yielding fast inference and an interpretable confidence score.
  4. A validation and ablation protocol: Multi-stage LLM-as-judge validation of both claim atomicity and triple correctness (Cohen's kappa = 0.892), plus a one-leave-out ablation study isolating the contribution of each augmentation module.

Main Findings

  • Large in-domain gains over the baseline: On FISCAL-data, MiniCheck-FISCAL (7B) reaches Precision 87.94, Recall 84.98, F1 86.43, and Accuracy 86.66, versus MiniCheck-7B's Precision 79.72, Recall 58.18, F1 67.27, and Accuracy 71.69. The paper reports Recall increasing by 26.8 points and Precision by 8.22.
  • Competitive with far larger open-weight models: On FISCAL-data, Mixtral-8x22B (141B) scores F1 89.66 and Accuracy 89.18, and C4AI Command R+ (104B) scores F1 87.82 and Accuracy 87.00 — close to the 7B MiniCheck-FISCAL's F1 86.43. Qwen2-72B scores F1 31.96 despite its size.
  • Outperforms GPT-3.5 Turbo in-domain, approaches GPT-4o: On FISCAL-data, GPT-3.5-turbo records F1 82.90 and Accuracy 83.58, while GPT-4o records F1 90.39 and Accuracy 89.52. Gemini-1.5-Flash reaches F1 88.56 and Claude-3.5-Sonnet F1 87.39.
  • Generalizes to external benchmarks: On FinDVer, MiniCheck-FISCAL improves F1 by +10.84 over MiniCheck-7B (70.53 vs. 59.69), with Accuracy 75.60 vs. 69.20. On Fin-Fact, it improves F1 by +7.55 (60.69 vs. 53.14), with Accuracy 62.59 vs. 58.61. Recall gains are consistent on both, while Precision stays high.
  • Strong on the FDV-IE subset under a RAG setting: MiniCheck-FISCAL (7B) achieves 75.60, exceeding same-size peers such as Mistral-7B-v3 (59.5), Gemma-7B (59.5), and Llama-2-7B (60.0), surpassing larger open models such as Qwen2-72B (68.0) and Mixtral-8x22B (70.0), and exceeding Gemini-1.5-Flash (70.5) while approaching GPT-4o (78.5) and GPT-3.5-turbo (79.0). Claude-3.5-Sonnet leads this subset at 80.5.
  • Ablation shows the modules are complementary: Removing Claim Paraphraser yields the highest precision (90.44) but collapses recall to 53.03 (F1 66.86). Removing Mis-attribution maximizes recall (88.68) but lowers precision to 82.48. Removing Summarization keeps precision high (89.97) but drops recall to 79.48 (F1 84.40). Removing Conflict Insertion lowers precision to 80.53 while recall stays stable at 84.87 (F1 82.64). Removing Fact Exclusion (F1 86.08) or Fact Distortion (F1 86.28) causes only modest declines, suggesting redundancy that bolsters robustness.
  • Efficiency is the central claim: The paper argues that at 7B parameters, MiniCheck-FISCAL offers lower cost, faster inference, and easier integration in compliance-sensitive environments than massive LLMs or closed APIs, with a single-token inference step and directly interpretable confidence scores.

Methodology in Plain English

  1. Start with real financial documents. The team draws on FinanceBench, which provides question–answer pairs with ground-truth evidence. The evidence contexts become the base documents, keeping the data realistic.
  2. Extract numerical claims. Because numbers are a primary source of hallucination in finance, a prompt-based pipeline using Qwen3 32B identifies quantitative statements. Each extracted claim is checked for "atomicity" — whether it expresses one single, checkable fact — using LLM judges.
  3. Create both supported and unsupported cases. For each claim–document pair, the modular synthesizer generates positive and negative examples. One module paraphrases the claim professionally while preserving all facts. Others alter the document: inserting a subtle contradiction, deleting the supporting evidence, distorting values or dates, reassigning attribution (year, entity, category), or summarizing while retaining only claim-relevant details. This spreads coverage across easy and hard error types.
  4. Validate with independent judges. Claims are reviewed by GPT-OSS-120B and Llama4 Maverick for atomicity, and only unanimous claims are kept. Triplets are validated by GPT-OSS-120B and Llama4 Scout, with agreement measured by Cohen's kappa at 0.892; only unanimously agreed triplets are preserved.
  5. Fine-tune a small verifier. MiniCheck-7B is adapted with LoRA. Instead of standard classification, verification is cast as causal language modeling: the model reads a fixed system instruction plus the claim–document pair and generates a single token, "yes" if the document supports the claim and "no" otherwise, trained on the negative log-likelihood of the correct token.
  6. Predict with a threshold. At inference, the model's probability for the "yes" token is used as a confidence score, and a threshold is applied to produce a binary decision — keeping inference cheap and the output interpretable.

Why This Matters

  • Impact on research: The paper argues that domain-specific synthetic data plus efficient fine-tuning can let compact models reach accuracy levels that would otherwise require brute-force scaling, positioning data quality and targeted augmentation as an alternative to ever-larger models. It also contributes an open dataset and evaluation setup for financial fact-checking.
  • Real-world applications:
    • Automated verification of numerical claims in financial filings, earnings reports, and 10-K disclosures.
    • Compliance and audit workflows where hallucinated figures carry regulatory or financial risk.
    • Retrieval-augmented question answering over financial documents, where a lightweight verifier can sit between retrieved evidence and a generated answer.
    • Investment and market analysis tooling that must flag unsupported quantitative statements quickly and cheaply.
  • Industry relevance: At 7B parameters, the model offers lower cost, faster inference, and easier integration in compliance-sensitive environments than 100B+ open-weight systems or closed proprietary APIs. The paper's emphasis on high recall matters because in financial fact-checking, missing a relevant claim can be as costly as misclassifying it.

Future Directions

  • Move beyond binary verification: The authors note that framing verification as a simple "yes"/"no" decision omits partial support and multi-hop evidence, and that richer task formulations and evaluation criteria are needed.
  • Close the synthetic-to-real gap: Synthetic perturbations cannot fully capture naturally occurring reporting errors such as subtle narrative shifts, multi-hop inconsistencies, or context-specific omissions.
  • Reduce reliance on LLM judges: Using LLMs as both generators and judges may introduce systematic biases relative to human annotators, even with multi-stage validation and unanimous-agreement filtering.
  • Add multimodal evidence and improve interpretability: The conclusion explicitly names multimodal evidence and results' interpretability as aims for future work.

Target Audience

  • Researchers working on fact verification, claim checking, or synthetic data generation for domain-specific NLP.
  • Financial NLP practitioners who need accurate, low-cost verification of numerical claims in filings and reports.
  • Applied machine learning engineers evaluating whether small, LoRA-tuned models can replace much larger systems in production.
  • Compliance, audit, and risk teams interested in tooling that produces fast, interpretable support/contradiction decisions on financial documents.

Authors’ abstract

Financial applications of large language models (LLMs) require factual reliability and computational efficiency, yet current systems often hallucinate details and depend on prohibitively large models. We propose FISCAL (Financial Synthetic Claim-Document Augmented Learning), a modular framework for generating synthetic data tailored to financial fact-checking. Using FISCAL, we generate a dataset called FISCAL-data and use it to train MiniCheck-FISCAL, a lightweight verifier for numerical financial claims. MiniCheck-FISCAL outperforms its baseline, surpasses GPT-3.5 Turbo and other open-source peers of similar size, and approaches the accuracy of much larger systems (20x), such as Mixtral-8x22B and Command R+. On external datasets FinDVer and Fin-Fact, it rivals GPT-4o and Claude-3.5 while outperforming Gemini-1.5 Flash. These results show that domain-specific synthetic data, combined with efficient fine-tuning, enables compact models to achieve state-of-the-art accuracy, robustness, and scalability for practical financial AI. The dataset and scripts are available in the project repository (link provided in the paper).

Read the original paper