Research
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
Overview Research area: Natural Language Processing — specifically the evaluation of large language models generating faith-sensitive Islamic content, combining Islamic NLP, retrieval-augmented/agenti
- arXiv
- 2510.24438
- Published
- 2025-10-28
- Authors
- Abdullah Mushtaq, Rafay Naeem, Ezieddin Elmahjub, Ibrahim Ghaznavi, Shawqi Al-Maliki, Mohamed Abdallah, Ala Al-Fuqaha, Junaid Qadir
AI summary
Overview
Research area: Natural Language Processing — specifically the evaluation of large language models generating faith-sensitive Islamic content, combining Islamic NLP, retrieval-augmented/agentic verification, and high-stakes domain evaluation.
Technical level: Intermediate. The framework is conceptually straightforward (prompting, tool use, rubric scoring), but readers will benefit from familiarity with LLM evaluation, RAG pipelines, and citation-verification concepts.
One-sentence scope: A pilot study that builds a dual-agent (quantitative plus qualitative) evaluation pipeline and applies it to GPT-4o, Ansari AI, and Fanar on 50 Islamic-content prompts to test whether these models can write theologically accurate, correctly cited, and respectfully toned Islamic essays.
What This Paper Is About
LLMs are increasingly used for Islamic guidance, but they risk misquoting Qur'anic verses, misattributing Hadiths, misapplying jurisprudence, or using inappropriate tone — errors with spiritual, cultural, and sometimes physical consequences. Conventional metrics such as BLEU or ROUGE only capture surface overlap and cannot judge doctrinal fidelity, citation integrity, or reverent style, and no existing pipeline unified theological verification with stylistic evaluation for Islamic writing. The authors build one, applying it to three models across 50 prompts drawn from authentic Islamic blog titles to test how faithfully current systems perform.
Key Contributions
- A dual-agent evaluation framework. A quantitative agent (OpenAI's o3 reasoning model) performs citation verification and 1–5 scoring across six criteria, while a qualitative agent performs side-by-side comparison of all three models' responses with justification-driven "Best"/"Worst" verdicts. Verification flags (confirmed / partially confirmed / unverified / refuted) are compiled into an
accuracy_verification_log. - A pilot benchmark dataset. 50 prompts collected from titles of blogs by recognized Islamic scholars across The Thinking Muslim, IslamOnline, Yaqeen Institute, SeekersGuidance, and UlumalHadith, spanning five domains: Jurisprudence (Fiqh), Qur'anic Exegesis (Tafsir), Hadith Sciences (Ulum al-Hadith), Theology (Aqidah), and Spiritual Conduct (Adab). These produced 150 essays archived verbatim as prompt–response pairs.
- The first systematic comparative evaluation of Islamically faithful generation across a general-purpose model (GPT-4o), an Islamic chatbot (Ansari AI), and a region-focused Arabic/Islamic model (Fanar), reported at both aggregate and category level.
- An explainable, reference-level audit trail. Worked examples trace individual Qur'anic references and external sources to show how the pipeline detects citation hallucinations and provides evidence-backed justifications — plus a public release of prompts, responses, and a complete code repository on GitHub.
Main Findings
- Overall quantitative ranking: GPT-4o achieved the highest overall mean score (3.90/5), followed by Ansari AI (3.79), with Fanar trailing at 3.04.
- Stability: GPT-4o showed the lowest response variability (std = 0.589); Fanar fluctuated most (std = 0.923).
- Islamic Accuracy: GPT-4o scored highest (3.93), Ansari AI close behind (3.68), Fanar lowest (2.76).
- Citation / Islamic Source Use: GPT-4o led (3.38), Ansari AI nearly matched it (3.32), and Fanar lagged substantially (1.82).
- Style & Structure: GPT-4o led in Theme (4.43) and Structure (4.16); Fanar performed worst in Originality (2.73).
- Variance in the hardest dimensions: Fanar showed the highest fluctuation in Islamic Accuracy (0.986) and Citation (0.727); GPT-4o and Ansari AI showed moderate inconsistency, meaning even strong models struggle to maintain citation fidelity across diverse prompts.
- Model design partially explains gaps: Fanar's smaller size (9B parameters) and limited context window (4,096 tokens) constrain nuanced Islamic reasoning and citation, while GPT-4o's larger context (128K tokens) supports stronger coherence, accuracy, and style. Fanar nonetheless contributes a morphology-based tokenizer, region-specific datasets, and an Islamic RAG pipeline.
- Qualitative hierarchy: Ansari AI earned 116 "Best" verdicts against only 3 "Worst"; GPT-4o earned 84 "Best" against 4 "Worst"; Fanar received no "Best" ratings and 193 "Worst" verdicts (50 in Clarity & Structure, 46 in Islamic Accuracy, 47 in Tone & Appropriateness, 50 in Depth & Originality).
- GPT-4o's qualitative strength was stylistic nuance: 48 "Best" in Tone & Appropriateness and 19 in Depth & Originality; Ansari AI's strengths were clarity and religious fidelity (41 "Best" in Clarity & Structure, 42 in Islamic Accuracy, 31 in Depth & Originality).
- Reporting note: The qualitative agent is described in the methodology as evaluating five dimensions (including Comparative Reflection), while the abstract and conclusion describe four writing dimensions and the reported verdict totals are out of 200.
- Citation failures persist at the top: Despite relatively strong scores, all models still fall short on reliable citation handling, faithful reference use, and contextual integrity.
- Case-level example (Appendix A.3): For a Fanar response, the agent confirmed a citation to Surah 49:13 was correctly quoted and contextually appropriate, but refuted a claim that verse 2:282 supports equal testimonial capacity (2:282 concerns financial testimony), detected a hallucinated link in which a reference to 2:282 pointed to verse 65:2 on Qur'an.com, and confirmed that 65:2 concerns divorce regulations and is unrelated to the topic. The agent summarized that only two Qur'anic references were provided, only one was accurate, with no primary hadith citations or verifiable external sources.
- Examples from the verification logs (Appendix A.4): A Fiqh claim that the intention to offer Udhiyyah must be made before the first day of Dhul-Hijjah was refuted (no classical source obliges this); a Tafsir response quoted "There is no compulsion in religion" as Qur'an 2:62 when the text matches 2:256 (wrong verse number); a Tafsir response quoted "To you your religion, to me mine" as Qur'an 29:46 when that verse does not contain the phrase (misquotation and mis-context); an IslamOnline Istikhārah link was valid and relevant but unused by the essay; a claimed Hadith about the Prophet beginning each day with the Basmala was unverified; and a Mawdudi quotation could not be located in commonly available editions of Towards Understanding Islam or Islamic Way of Life. Partially confirmed cases included a downplaying of IslamQA's treatment of birthday celebrations (#1027), an incomplete statement that Udhiyyah is a confirmed Sunnah (Hanafis deem it wajib, per IslamWeb article 171933), and a non-verbatim paraphrase of Sahih al-Bukhari 1166 presented without a reference.
- Category-level detail (Table 1): GPT-4o's best category performance came in Spiritual Conduct (Adab), with Theme 4.75 and Structure 4.40; Fanar's lowest values appeared in Citation across categories (1.55 in Jurisprudence, 1.65 in Tafsir, 1.80 in Aqidah, 2.20 in Hadith, 1.90 in Adab).
- Human sanity check: A human evaluator reviewed the agent's outputs, and no modifications were deemed necessary, though the review was used to confirm alignment with evaluation objectives and to flag areas for future refinement.
Methodology in Plain English
The researchers first wrote 50 prompts by taking the titles of blogs written by recognized Islamic scholars on reputable platforms, covering five subject areas so the test would not be thematically narrow. Each prompt was wrapped in a fixed template asking for a thorough, clear, blog-style essay aimed at a general audience. The same prompts were sent to ChatGPT (GPT-4o), Ansari AI, and Fanar, and the 150 resulting essays were stored word-for-word.
Evaluation happens in two passes. The quantitative agent, built on OpenAI's o3 reasoning model, has access to three verification tools (Qur'an Ayah, Internet Search, Internet Extract). It splits each essay into introduction, body, and conclusion and scores it from 1 to 5 on six criteria: Structural Coherence, Thematic Focus, Clarity, Originality, Islamic Accuracy, and Citation/Islamic Source Use. The first four roll up into a "Style and Content Evaluation" dimension and the last two into an "Islamic Content Evaluation" dimension. When the agent spots a reference, the tools fetch the relevant verse, Hadith, or source text and return a flag — confirmed, partially confirmed, unverified, or refuted — with points deducted for the weaker flags.
The qualitative agent puts all three models' answers side by side for each prompt, separated by <R1>, <R2>, and <R3> tags, and judges them across five described dimensions (Clarity & Structure, Islamic Accuracy, Tone & Appropriateness, Depth & Originality, Comparative Reflection). For each dimension it names the strongest and weakest response, justifies the choice with exact text excerpts, and checks religious content with the same toolchain. Aligning the qualitative judgments with the quantitative scores provides early evidence of convergent validity. A human evaluator then reviewed the agent's outputs as a sanity check.
Why This Matters
Impact on research. The paper argues that religious content generation is a high-stakes domain as consequential as medicine, law, and journalism, yet lacks the domain-specific evaluation pipelines those fields already have. It offers a modular, interpretable template that links generated text to reference-level verification, so evaluation is not just a number but an auditable trail of why each citation was accepted or rejected. It also argues that standard metrics and even general Arabic benchmarks (Arabic-SQuAD, MLQA, TyDiQA, Arabic MMLU) mostly test linguistic competence rather than theological grounding, and that under-digitized, unstructured, fragmented Islamic corpora remain a structural barrier to progress.
Real-world applications.
- Islamic chatbots and fatwa-adjacent assistants could use the pipeline as a pre-deployment gate, catching misquoted verses and misattributed Hadiths before users see them.
- Mosques, Islamic educational platforms, and publishers considering AI-assisted content could require citation-verification logs before anything is published.
- Scholars and madrasa instructors could use the verdict tables and evidence excerpts as a teaching tool about where models fail and why.
- The same architecture could transfer to other faith-sensitive or high-stakes writing, such as biblical citation in theological education, medical guidance, or legal drafting.
Industry relevance. The paper situates itself against documented failures elsewhere: general chatbots hallucinate at 58–82% on legal questions, RAG-backed legal tools still err at over 17% (Lexis+ AI, Thomson Reuters' Practical Law) and over 34% (Westlaw), SourceCheckup found 50–90% of medical responses not fully supported by their own citations with GPT-4 plus RAG still at 30% unsupported statements, and CNET corrected 41 of 77 AI-written finance articles. For vendors deploying faith-based or expert-facing assistants, the implication is that fluent output is not sufficient evidence of reliability and that disclaimer, scholar oversight, and verifiable sourcing must be built into the product rather than assumed.
Future Directions
- Reduce evaluator bias through architectural diversity. The blind protocol mitigates within-family bias, but future work should use a heterogeneous ensemble of evaluator LLMs (for example Claude, Gemini, and Llama) for cross-validation and report inter-evaluator agreement across model families.
- Scale and multilingual validation. Expand beyond the 50 prompts with stratified sampling across madhahib, edge cases, and both classical and contemporary jurisprudence; build parallel Arabic-first evaluations with native-speaking scholars, test cross-lingual consistency, and evaluate Arabic-targeted systems such as Fanar in their primary language.
- Multi-expert human validation. Convene panels of 3 to 5 Islamic scholars per prompt with diversity across madhab, geography, and specialization, then measure consensus and adjudicate disagreements.
- Establish responsible-use standards. Current models fall short on faith-sensitive rigor and citation integrity, so the authors call for clear disclaimers, mandatory scholar oversight, and community-driven evaluation that reflects diverse Islamic perspectives — positioning AI to assist rather than replace human religious scholarship.
Target Audience
Researchers and practitioners working on Islamic NLP, Arabic LLMs, and faith-sensitive content generation; AI safety and evaluation scientists interested in citation grounding and agentic verification in high-stakes domains; Islamic scholars and institutional reviewers who need to judge whether a model's output is theologically acceptable; and product teams building religious, medical, legal, or journalistic assistants where misattributed sourcing carries real consequences.
Authors’ abstract
Large language models are increasingly used for Islamic guidance, but risk misquoting texts, misapplying jurisprudence, or producing culturally inconsistent responses. We pilot an evaluation of GPT-4o, Ansari AI, and Fanar on prompts from authentic Islamic blogs. Our dual-agent framework uses a quantitative agent for citation verification and six-dimensional scoring (e.g., Structure, Islamic Consistency, Citations) and a qualitative agent for five-dimensional side-by-side comparison (e.g., Tone, Depth, Originality). GPT-4o scored highest in Islamic Accuracy (3.93) and Citation (3.38), Ansari AI followed (3.68, 3.32), and Fanar lagged (2.76, 1.82). Despite relatively strong performance, models still fall short in reliably producing accurate Islamic content and citations -- a paramount requirement in faith-sensitive writing. GPT-4o had the highest mean quantitative score (3.90/5), while Ansari AI led qualitative pairwise wins (116/200). Fanar, though trailing, introduces innovations for Islamic and Arabic contexts. This study underscores the need for community-driven benchmarks centering Muslim perspectives, offering an early step toward more reliable AI in Islamic knowledge and other high-stakes domains such as medicine, law, and journalism.