Skip to content
AI.info

Research

VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation

Overview Research area: Ontology engineering and neuro-symbolic AI — specifically LLM-based generation of Competency Questions (CQs) for validating OWL ontologies. Technical level: Advanced. The paper

VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation
arXiv
2511.07991
Published
2025-11-11
Authors
Hyojun Choi, Seokju Hwang, Kyong-Ho Lee

AI summary

Overview

Research area: Ontology engineering and neuro-symbolic AI — specifically LLM-based generation of Competency Questions (CQs) for validating OWL ontologies.

Technical level: Advanced. The paper assumes familiarity with description logics, OWL axioms (e.g., allValuesFrom, someValuesFrom, union/intersection), TBox validation, and LLM fine-tuning with LoRA.

Scope: The paper introduces VSPO, a dataset and fine-tuned LLaMA-3.1-8B-Instruct model that generates competency questions designed to expose "semantic pitfalls" — mismatches between an ontology term's natural language definition and its formal axioms that reasoners cannot detect.

What This Paper Is About

Ontology engineers write competency questions (CQs) to check whether an ontology actually encodes the knowledge it is supposed to. Writing these by hand is slow and expensive, so recent work has used large language models to generate them automatically — but those approaches judge success by cosine similarity to existing question sets, which does not test whether a generated question can catch logical modeling errors. This paper builds a training dataset and a fine-tuned model that generate CQs aimed specifically at a class of errors the authors call semantic pitfalls: cases such as "Misusing allValuesFrom" where the natural language definition and the formal axiom quietly disagree, and standard reasoners cannot flag the problem.

Key Contributions

  1. A new perspective on LLM-based CQ generation: evaluation shifts from similarity to existing human-authored CQs toward the model's ability to detect semantic pitfalls in an ontology.
  2. VSPO, a dataset of 1,563 instances built from six ontologies by injecting controlled misalignments between generated natural language definitions and ontology axioms, split into 1,368 training and 195 test instances.
  3. A taxonomy of four misalignment types — Type 1 (missing axiom), Type 2 (undefined axiom), Type 3 (misusing axiom), and Type 4 (alignment) — where Types 1–3 deliberately create semantic pitfalls and Type 4 preserves consistency.
  4. A fine-tuned LLaMA-3.1-8B-Instruct model that reports 26% higher precision and 28.2% higher recall than GPT-4.1 in generating CQs for pitfall validation, described by the authors as the first study to target semantic pitfall validation in CQ generation using LLMs.

Main Findings

  • VSPO outperforms both baselines on pitfall-focused CQs (CQ_sp): VSPO reached precision 75.0, recall 62.2, F1 68.0, and average maximum cosine similarity 0.7950. GPT-4.1 reached 49.0 / 34.0 / 40.1 / 0.6588, and LLaMA-3.1-8B-Instruct reached 29.8 / 20.5 / 24.3 / 0.5927.
  • VSPO also leads on general axiom-based CQs (CQ_normal): precision 95.9, recall 35.8, F1 52.1, C.S. 0.8708, versus GPT-4.1 at 82.1 / 27.1 / 40.7 / 0.7871 and LLaMA-3.1-8B-Instruct at 58.5 / 20.2 / 30.0 / 0.7156. The paper notes precision is the more meaningful metric here because CQ_sp is a subset of CQ_normal and the larger reference set lowers recall.
  • The advantage holds across every misalignment type: VSPO scored highest in all metrics for Type 1 (71.1 / 54.4 / 61.6 / 0.7867), Type 2 (83.8 / 69.4 / 75.9 / 0.8172), and Type 3 (69.0 / 63.2 / 66.0 / 0.7777). Its precision on Type 2 (83.8) and Type 3 (69.0) were described as significantly exceeding the baselines.
  • Strong performance even when no pitfall exists: In the Type 4 alignment setting, VSPO achieved 96.7 precision, 68.8 recall, F1 80.4, and C.S. 0.8722, indicating it still generates accurate CQs without inconsistencies to exploit.
  • Generalization to unseen ontologies: Under a leave-one-out setting across AWO, DEM@Care, Stuff, SWO, OntoDT, and Pizza, VSPO scored F1 between 53.0 (Pizza) and 68.5 (OntoDT). The paper attributes the weaker Pizza result to its disproportionately high number of disjointWith and subClassOf axioms.
  • Threshold sensitivity is a flaw in prior evaluation practice: The authors observe that LLaMA-3.1-8B-Instruct's maximum cosine similarity stayed similar across the three pitfall types while precision, recall, and F1 fluctuated sharply, which they attribute to sensitivity to the chosen similarity threshold.
  • Qualitative case study on the herbivore axiom: Given the axiom [herbivore EquivalentTo ((eats only plant) and (eats only (is-part-of some plant)))], where a union should have been used, VSPO asked "Does being a herbivore entail eating only things that are either directly a plant or part of a plant?", while GPT-4.1 asked generic questions such as "Which animals are herbivores?" and LLaMA-3.1-8B-Instruct asked "What do herbivores eat?" — neither targeting the intersection/union confusion.

Methodology in Plain English

The researchers start from six existing ontologies: AWO, DEM@Care, SWO, Stuff, OntoDT, and Pizza. For SWO, which contains 3,993 classes and 56 properties, they randomly sampled 500 ontology terms; for the rest they used all classes and properties. Total term counts per ontology were 32 (AWO), 411 (Dem@Care), 500 (SWO), 94 (Stuff), 419 (OntoDT), and 107 (Pizza).

For each ontology term, they extract the axioms in which it is the subject and then run two parallel processes. First, a template-based generator produces candidate competency questions, using prompt templates written by hand for each axiom type (3 to 7 templates per type) with GPT-4.1, generating n = 3 CQs per axiom. Second, a type classifier randomly assigns each term to one of four misalignment categories, provided the term has the axioms that category requires (two or more axioms for Types 1 and 2; someValuesFrom/allValuesFrom or intersection/union constructs for Type 3). Terms that qualify for nothing are assigned only to Type 4.

The misalignment injection then alters the picture: in Type 1 a random axiom is deleted from the input axiom set while the definition is written from the complete set; in Type 2 the axiom set stays complete but one axiom is withheld when generating the definition; in Type 3 logical operators such as someValuesFrom/allValuesFrom or intersection/union are swapped inside one axiom while the definition uses the original axioms; in Type 4 nothing changes. The CQ generated for the disrupted axiom becomes the training target, called CQ_sp.

The resulting (A_T, D_T, CQ_sp) triples fine-tune LLaMA-3.1-8B-Instruct with LoRA (rank r = 8, scaling factor α = 16, dropout 0.05), 3 epochs, effective batch size 4, learning rate 3 × 10⁻⁴, and bf16 precision, on two NVIDIA RTX 3090 GPUs. The paper reports that training beyond 3 epochs caused overfitting. Evaluation uses SentenceBERT cosine similarity with a validity threshold τ = 0.7, plus the average maximum cosine similarity per generated question.

Why This Matters

Impact on research. The paper reframes what a "good" generated competency question is. Prior LLM-based CQ work (including studies by Rebboud et al., Alharbi et al., and Pan et al.) scored outputs by cosine similarity to existing CQs; VSPO scores them by whether they expose a logical inconsistency. The authors also document that fixed-threshold similarity metrics are unstable, which is a methodological caution for the field.

Real-world applications.

  • Auditing biomedical and clinical ontologies such as SWO and Dem@Care, where a silent axiom error could distort downstream reasoning in patient-care or software-annotation systems.
  • Quality control in large public ontology repositories, complementing rule-based scanners like OOPS! (which the paper notes was built by empirically analyzing over 693 ontologies).
  • Educational tooling for ontology engineering courses, using ontologies like Pizza and AWO to show students where definitions and axioms diverge.
  • Supply-chain and materials modeling contexts such as the Stuff materials ontology, where subtle logical restrictions affect classification.

Industry relevance. The paper argues that generating sophisticated CQs with open-source models matters especially for sensitive or security-critical knowledge domains. It also positions the work as reducing the human bottleneck in building symbolic knowledge that LLMs do not inherently possess.

Future Directions

  • Extending the framework to generate not just CQs but their corresponding SPARQL-OWL queries, possibly by generating the query first and deriving the natural language question from it.
  • Addressing the limitation that misalignments were injected one axiom at a time, since cases with multiple simultaneous semantic pitfalls were not covered.
  • Expanding beyond the six ontologies and correcting the imbalanced axiom type distribution — the paper notes rdf:type properties were extremely rare, with only 11 instances in total.
  • Closing the gap between the highly refined generated definitions and real-world ontologies, whose natural language annotations are typically sparse or ambiguous, and moving toward a fuller end-to-end automated validation pipeline, since VSPO still requires human verification of its output.

Target Audience

Ontology engineers and knowledge representation researchers who need to validate TBox designs; NLP researchers working on LLM-based generation of competency questions; practitioners in biomedical, clinical, and materials informatics who maintain OWL ontologies; and readers interested in neuro-symbolic methods where language models assist rather than replace formal reasoning.

Authors’ abstract

Competency Questions (CQs) play a crucial role in validating ontology design. While manually crafting CQs can be highly time-consuming and costly for ontology engineers, recent studies have explored the use of large language models (LLMs) to automate this process. However, prior approaches have largely evaluated generated CQs based on their similarity to existing datasets, which often fail to verify semantic pitfalls such as "Misusing allValuesFrom". Since such pitfalls cannot be reliably detected through rule-based methods, we propose a novel dataset and model of Validating Semantic Pitfalls in Ontology (VSPO) for CQ generation specifically designed to verify the semantic pitfalls. To simulate missing and misused axioms, we use LLMs to generate natural language definitions of classes and properties and introduce misalignments between the definitions and the ontology by removing axioms or altering logical operators (e.g., substituting union with intersection). We then fine-tune LLaMA-3.1-8B-Instruct to generate CQs that validate these semantic discrepancies between the provided definitions and the corresponding axioms. The resulting CQs can detect a broader range of modeling errors compared to existing public datasets. Our fine-tuned model demonstrates superior performance over baselines, showing 26% higher precision and 28.2% higher recall than GPT-4.1 in generating CQs for pitfall validation. This research enables automatic generation of TBox-validating CQs using LLMs, significantly reducing manual effort while improving semantic alignment between ontologies and expert knowledge. To the best of our knowledge, this is the first study to target semantic pitfall validation in CQ generation using LLMs.

Read the original paper