Skip to content
AI.info

Research

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

Overview Research area: Multimodal AI safety — jailbreak detection for Vision Language Models (VLMs), specialized to the financial domain. Submitted under arXiv category Natural Language Processing (c

FENCE: A Financial and Multimodal Jailbreak Detection Dataset
arXiv
2602.18154
Published
2026-02-20
Authors
Mirae Kim, Seonghun Jeong, Youngjun Kwak

AI summary

Overview

  • Research area: Multimodal AI safety — jailbreak detection for Vision Language Models (VLMs), specialized to the financial domain. Submitted under arXiv category Natural Language Processing (cs.CL), arXiv:2602.18154v3.
  • Technical level: Intermediate. The paper is a dataset-and-benchmark paper; the core ideas are accessible, but it assumes familiarity with VLMs, attack success rate, fine-tuning, and standard classification metrics.
  • Scope in one sentence: This paper introduces FENCE, a bilingual (Korean–English) 10k-sample finance-domain text–image dataset built to measure jailbreak vulnerability in VLMs and to train guardrail detectors for financial applications.

What This Paper Is About

Jailbreaking — manipulating a model into producing harmful or unintended responses — is a known risk for large language models, and VLMs are especially exposed because they process both text and images, widening the attack surface. Existing jailbreak datasets are mostly general-purpose, English-centered, and focused on evaluation rather than training defenses, leaving the financial sector, which handles sensitive data under strict regulation, largely uncovered. FENCE addresses this gap by supplying a finance-specific, image-based, bilingual dataset for both diagnosing VLM vulnerabilities and training detectors.

Key Contributions

  1. A finance-focused multimodal jailbreak dataset. FENCE contains 10k text–image pairs (5,000 benign and 5,000 harmful, a balanced 50:50 ratio) spanning more than 15 finance categories such as loans, deposits, credit cards, and online banking. It is described by the authors as the first benchmark explicitly designed for jailbreak evaluation and mitigation in finance-focused multimodal systems.
  2. Emphasis on image-based attacks (IA). Unlike predominantly text-focused resources, FENCE is a fully IA dataset with three structural settings (BaseImg, TextImg, FigStep) and multiple layout strategies. The paper notes JailBreakV-28K contains only 28.6% IA data, whereas FigStep and MM-SafetyBench rely on a single fixed generation strategy.
  3. Bilingual construction. The dataset was built natively in Korean from a base of 2,500 unique benign Korean queries and then translated into English, yielding 5,000 Korean and 5,000 English instances so that Korean financial and linguistic nuance is preserved while remaining accessible to the broader research community.
  4. A trainable detector baseline. Fine-tuning lightweight PaliGemma and Qwen-family models on FENCE produces a binary harmful-query classifier reported at 99% in-distribution accuracy, with strong out-of-distribution behavior on external benchmarks and a large gain in defense success rate.

Main Findings

  • FENCE yields the highest overall attack success rate. Across the evaluated models, mean ASR was 17.76% (±13.50) on FENCE versus 15.74% (±18.77) on JailBreakV-28K, 16.31% (±17.28) on FigStep, 8.77% (±11.34) on HADES, and 12.21% (±11.82) on MM-SafetyBench.
  • Even well-aligned proprietary models are affected. GPT-4o recorded 4.60% ASR on FENCE compared with 0.00% (JailBreakV-28K), 0.20% (FigStep), 0.00% (HADES), and 1.40% (MM-SafetyBench). GPT-4o-mini recorded 12.20% on FENCE versus 0.71%, 12.40%, 0.00%, and 3.60% on the same four benchmarks.
  • Smaller open-source models are the most exposed. On FENCE, Qwen2.5-VL Instruct 3B reached 48.60% ASR, VARCO Vision 1.7B reached 40.80%, Qwen2.5-VL Instruct 32B reached 29.80%, and Kanana1.5 Vision Instruct 3B reached 27.40%.
  • Higher ASR is not framed as weaker safety. The authors argue FENCE targets finance-specific attack surfaces that general-purpose safety training does not cover, exposing latent vulnerabilities existing benchmarks miss.
  • FigStep-style typographic overlays are the strongest attack type. Averaged over models, FigStep produced the highest ASR at 28.4% for Korean and 17.2% for English. English BaseImg was the least effective vector overall, at a mean ASR of 4.1%.
  • Korean inputs are generally less well guarded. Korean queries elicited higher ASRs across attack types (28.4% versus 17.2% for FigStep). The pattern reverses for Korean-specialized models: VARCO Vision 1.7B showed higher English ASR for TextImg (54.0% versus 25.0%) and FigStep (56.0% versus 48.0%), and Kanana1.5 Vision 3B showed the same reversal for FigStep (45.0% versus 28.0%).
  • Image-text recognition quality tracks detection performance. PaliGemma2 achieved the strongest instruction-based ITR results (English: EMR_inst 92.2%, STS_inst 0.90; Korean: 77.0%, 0.90) and the highest classification scores (FENCE accuracy 0.98 English / 0.97 Korean; JailBreakV-28K 0.78). PaliGemma1 was weaker (English EMR_inst 75.9%, STS_inst 0.76; Korean 58.7%, 0.77), with FENCE accuracy 0.94 and JailBreakV-28K accuracy 0.50.
  • Domain shift costs performance but does not break it. The FENCE-trained PaliGemma classifier reached 94–98% on its native test split, dropping to 78% on the mini subset of JailBreakV-28K, which the authors attribute to that benchmark being mostly English, text-only, and only loosely finance-related.
  • A compact FENCE-tuned model beats large guardrails. A Qwen2.5-VL 3B fine-tuned on FENCE reached a mean score of 0.99 across the five benchmarks, compared with 0.51 for LlamaGuard 3 Vision and 0.69 for LlamaGuard 4 in Table 6. (The Table 6 columns list sizes of 8B and 11B, while the surrounding text describes these baselines as LlamaGuard3 Vision (11B) and LlamaGuard4 (12B).)
  • Fine-tuning sharply improves rejection of harmful queries. Mean Defense Success Rate for Qwen2.5-VL 3B rose from 66.69% to 99.34% after fine-tuning, a gain of 32.65 percentage points. The largest gains were on FENCE (51.40% to 99.60%, +48.20) and FigStep (62.00% to 100.00%, +38.00); JailBreakV-28K went from 79.64% to 99.29% (+19.65), HADES from 68.40% to 99.60% (+31.20), and MM-SafetyBench from 72.00% to 98.20% (+26.20). Section 4.3 text states an average of 66.29% before fine-tuning, while Table 7 lists 66.69%.
  • Labels were validated by humans. GPT-4o was used as the automated evaluator, and human verification on 250 harmful Korean queries — 10% of the generated set — reached 95% agreement with GPT-4o's judgments.
  • FENCE's harmful queries are harder to separate from benign ones. A t-SNE visualization using text-embedding-3-small embeddings shows FENCE's harmful queries clustering closer to benign queries than those of existing datasets, indicating a more challenging classification setting.

Methodology in Plain English

The dataset was built in three stages.

First, the team collected real-world financial queries from the FAQs of six major South Korean financial institutions, producing 2,500 unique benign Korean samples. Rather than reuse existing harmful prompts, they transformed each benign query into a harmful counterpart using GPT-4o through a two-step prompting scheme: a role-playing prompt that reframes the question as something a financial criminal would ask while writing a crime-investigation drama script, followed by an evaluation prompt that judges whether the rewritten question is benign (b) or malicious (m). If judged benign, the transformation was retried for up to five attempts, after which it was treated as a failure. All benign–harmful pairs were then translated into English, giving 5,000 Korean and 5,000 English instances while preserving one-to-one semantic alignment.

Second, query-relevant images were gathered by keyword-based crawling of real-world photographs from Pixabay under its Content License — chosen over diffusion-generated images to avoid synthetic artifacts. Two image types were collected: semantically harm-aligned images manually curated for harmful queries, and neutral finance-related background images used as canvases.

Third, text and images were fused with standard image processing into three sample types: BaseImg (20% of samples), where an unmodified textual query is paired with a context-relevant image; TextImg (40%), where the query is overlaid onto a neutral finance background so the label depends only on the overlaid text; and FigStep (40%), where queries are rendered through FigStep layout templates that reorganize content into structured formats such as Question–Answer or Goal–Method.

For evaluation, FENCE was split 8:1:1 into 8,000 training, 1,000 validation, and 1,000 test samples. The test set retains the balanced distribution of 500 benign and 500 harmful queries; the 500 harmful test queries comprise 250 human-verified Korean queries and 250 English queries, and all 1,000 test samples were subsequently reviewed by human annotators. Two evaluation settings were used: in-distribution testing on the FENCE test split and out-of-distribution testing on four external benchmarks, using official mini versions where available. The primary metrics are Attack Success Rate and its complement Defense Success Rate, computed on harmful queries only, plus F1-score on the full balanced test set for classification experiments. A label of 1 (harmful) is assigned when either the text or the image component is harmful.

For the detection experiments, the authors focused on lightweight PaliGemma and Qwen architectures. They measured image-text recognition with Exact Match Ratio (EMR) and Semantic Textual Similarity (STS), using instruction-based variants in which GPT-4o replaces traditional cosine similarity as the scoring function, to check how recognition quality relates to classification accuracy.

Why This Matters

This work shifts multimodal safety research toward a high-stakes domain that existing general-purpose benchmarks only partially cover. It also provides a training corpus rather than just a test set: most jailbreak datasets contain only harmful samples, whereas FENCE's balanced benign/harmful design supports the discriminative training that guardrail models need. The finding that a finance-specific dataset can elicit harmful behavior from heavily aligned commercial models, and that fine-tuning a 3B model on it beats much larger guardrail baselines, has direct implications for how safety training data is curated and scoped.

Real-world applications:

  • Guardrails for financial chatbots and banking assistants. The dataset can be used to fine-tune or evaluate detectors that screen user inputs before a VLM responds, in a domain where leaked information or misleading output can cause fraud or privacy breaches.
  • Fraud and abuse prevention. Because harmful queries were derived from crime-planning framings around banking apps, credit reports, and accounts, the data reflects fraud-adjacent prompts that consumer-facing systems actually encounter.
  • Regulatory and compliance review. Financial institutions operate under strict regulation and rely on sensitive data; a domain-aware safety benchmark supports documenting and testing model safeguards before deployment.
  • Multilingual safety testing. The Korean–English split exposes language-specific gaps in alignment, useful for institutions serving non-English-speaking customers.

Industry relevance: the authors motivate the work by noting that South Korea has over 169 million mobile banking accounts, illustrating deep integration of AI into finance, and the dataset is released by Kakaobank, a commercial bank, through a public repository at https://github.com/kakaobank/FENCE.

Future Directions

  • Broaden language and scenario coverage. The limitations section states that FENCE's scale, domain scope, and Korean–English bilingual focus may limit broader linguistic generalization; future work plans additional languages and financial scenarios.
  • Add human-authored adversarial examples. Because FENCE is built from synthetic adversarial prompts generated by GPT-4o, the authors note it may not capture the full variety of real-world user behavior and plan to incorporate human-written attacks for greater realism.
  • Integrate FENCE into safety-tuning workflows. The authors intend to fold the dataset into training pipelines for robust multimodal models in financial applications.
  • Validate across more models and training paradigms. Evaluation covered a limited number of commercial and open-source VLMs, leaving room for testing more model families. A related open question is how to define "harm" in financial contexts, which the authors describe as inherently complex and shaped by legal, regulatory, and institutional factors, potentially introducing subjectivity into annotation.

Target Audience

Researchers and engineers working on multimodal AI safety,

Authors’ abstract

Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE emphasizes domain realism through finance-relevant queries paired with image-grounded threats. Experiments with commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99 percent in-distribution accuracy and maintains strong performance on external benchmarks, underscoring the dataset's robustness for training reliable detection models. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains. Warning: This paper includes example data that may be offensive.

Read the original paper