Research
Ontolearn-A Framework for Large-scale OWL Class Expression Learning in Python
Overview Research area: Neuro-symbolic machine learning and knowledge representation — specifically OWL class expression learning (concept learning) in description logics over RDF knowledge graphs. Te

- arXiv
- 2510.11561
- Published
- 2025-10-13
- Authors
- Caglar Demir, Alkid Baci, N'Dah Jean Kouagou, Leonie Nora Sieger, Stefan Heindorf, Simon Bin, Lukas Blübaum, Alexander Bigerl, Axel-Cyrille Ngonga Ngomo
AI summary
Overview
- Research area: Neuro-symbolic machine learning and knowledge representation — specifically OWL class expression learning (concept learning) in description logics over RDF knowledge graphs.
- Technical level: Intermediate. Readers need some familiarity with OWL ontologies, description logics, and knowledge graphs, but the paper is a framework/tool description rather than a theoretical or empirical study.
- Scope in one sentence: The paper introduces Ontolearn, an open-source MIT-licensed Python library that bundles nine symbolic, neuro-symbolic and deep learning algorithms for learning OWL class expressions over large knowledge graphs, and adds LLM-based verbalization and SPARQL/triplestore support.
What This Paper Is About
Given a knowledge base plus sets of positive and negative examples, class expression learning aims to find a description-logic class expression that the positive examples are instances of and the negative examples are not. The paper argues that most symbolic learners cannot operate well on knowledge graphs with millions of triples, and that the previously most mature system, DL-Learner, has not been actively maintained since its last release in 2021 and therefore lacks the latest neuro-symbolic models. Ontolearn is presented as the answer: a maintained, well-tested Python framework that packages recent learners together with reasoning, verbalization and remote-triplestore access.
Key Contributions
- A unified Python framework with nine learners. Ontolearn provides nine recent state-of-the-art symbolic, neuro-symbolic and deep learning algorithms for OWL class expression learning — EvoLearner, CLIP, NCES, NCES2, DRILL, Nero and ROCES — plus efficient Python implementations of CELOE and OCEL from DL-Learner. At the time of writing, the authors state it offers more OWL class expression learning algorithms than any other publicly available framework.
- Scalability beyond memory. Ontologies can be read into memory in RDF/XML, OWL/XML and N-Triples formats, or loaded into a triplestore such as Tentris when they exceed roughly 10^8 triples; OWL class expressions are then mapped to SPARQL queries for efficient instance retrieval.
- LLM-based verbalization. A verbalization module translates complex OWL class expressions into natural language sentences using large language models such as Llama or Mistral, illustrated on the "Married Female" example, where Female ⊓ (∃ married.⊤) becomes "A female who is married".
- Engineering maturity and ecosystem. The project comprises approximately 20,000 lines of code, ships 156 unit and regression tests with 95% test coverage, 26 example scripts and documentation, has been downloaded more than 26,000 times, and is installable via PyPI under the MIT license.
Main Findings
- Framework scope: Ontolearn supplies the knowledge base class, a reasoner component (Hermit and Pellet, implemented via Owlapy), the learning problem class, the concept learners, optional sampling, triplestore support, and the LLM verbalization module.
- DL-Learner wrappers included: Beyond reimplementing CELOE and OCEL in Python, Ontolearn provides wrappers giving access to the original DL-Learner implementations of those symbolic learners.
- No empirical benchmark results are reported in the paper. The paper describes architecture, features, usage and ecosystem; it does not report accuracy, F1, runtime or other comparative measurements for the learners.
- Test and quality metrics: 156 unit and regression tests, 95% test coverage, approximately 20,000 lines of code, and 26 example scripts are reported.
- Adoption metric: The library has been downloaded more than 26,000 times, according to the authors.
- Deployment options: It supports local ontology files or triplestore endpoints, includes the code needed to launch it as a web service, and provides helper functions for extending it with novel concept learners.
- Sampling support: Ontolearn supports sampling techniques to accelerate the learning process and triplestore implementations aimed at distributed and cloud-based environments.
Methodology in Plain English
Rather than proposing a new learning algorithm, the authors built an integration and engineering layer. They took existing class expression learners from the literature and implemented or wrapped them behind a common Python interface. A shared knowledge base component handles ontology loading, either into memory from standard OWL serialization formats or into a triplestore when the ontology is too large. A reasoner component, accessed through the Owlapy library, infers new knowledge and retrieves instances of candidate class expressions. A learning problem object bundles the knowledge base with positive and negative examples, which the learners consume. Learned expressions can be rendered as DL syntax or converted into SPARQL to run against a remote triplestore, and can be passed to an LLM to produce a natural-language sentence. Correctness is maintained through a test suite of 156 cases in Python's unittest framework.
Why This Matters
Impact on research. Ontolearn gives researchers a common, maintained codebase in which neuro-symbolic concept learners can be compared and reused, addressing the gap left by DL-Learner's inactivity since its 2021 release. Ante-hoc explainability — producing human-readable class expressions rather than opaque predictions — is positioned as a building block of trustworthy Web-scale AI, where RDF knowledge graphs are increasingly large.
Real-world applications reported in the paper:
- Industry 4.0 skill matching: Ontolearn has been used to automatically learn human-interpretable descriptions of machine skills, which the authors describe as crucial for tasks such as skill matching in production processes.
- Explaining graph neural networks: EDGE employs Ontolearn to explain graph neural networks in terms of class expressions.
- Feature selection and hyperparameter tuning: AutoCL facilitates feature selection and hyperparameter tuning for Ontolearn's concept learners.
- Preprocessing tabular data and ontologies: OntoSample applies graph sampling before the ontology is fed to Ontolearn for speedups while maintaining high predictive performance, and Tab2Onto automatically converts tabular data into an OWL ontology, easing application to industrial use cases.
Industry relevance. The framework is delivered as an installable PyPI package under a permissive MIT license, can run as a web service, and can query remote triplestores — properties that matter for deployments where knowledge does not fit in memory and where model outputs must be interpretable to domain experts.
Future Directions
- Empirical benchmarking. The paper reports no accuracy, F1 or runtime comparisons; systematic evaluation of the nine bundled learners on shared benchmarks is the obvious next step.
- Use of sampling for scale. The paper mentions sampling techniques only as an included feature; how much they speed up learning on very large graphs while preserving predictive performance remains an open question here.
- Verbalization quality. The LLM module is demonstrated on a single "Married Female" example; the fidelity of LLM-generated sentences for complex expressions is not evaluated in the paper.
- Ecosystem growth. The framework is explicitly designed to be extended with novel concept learners, and complementary tools such as EDGE, AutoCL, OntoSample and Tab2Onto already build on it — inviting further integration and new downstream applications.
Target Audience
This paper benefits machine learning and semantic web researchers who need explainable models over RDF knowledge graphs, knowledge engineers and ontology practitioners who want to apply concept learning without reimplementing algorithms, and industry developers in settings such as Industry 4.0 where interpretable, machine-learned descriptions of entities are required. It is also useful to developers seeking a maintained Python alternative to DL-Learner.
Authors’ abstract
In this paper, we present Ontolearn-a framework for learning OWL class expressions over large knowledge graphs. Ontolearn contains efficient implementations of recent stateof-the-art symbolic and neuro-symbolic class expression learners including EvoLearner and DRILL. A learned OWL class expression can be used to classify instances in the knowledge graph. Furthermore, Ontolearn integrates a verbalization module based on an LLM to translate complex OWL class expressions into natural language sentences. By mapping OWL class expressions into respective SPARQL queries, Ontolearn can be easily used to operate over a remote triplestore. The source code of Ontolearn is available at https://github.com/dice-group/Ontolearn.