Research
Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
Overview Research area: Natural Language Processing — relation extraction from text into structured knowledge graphs, with a focus on small language models (SLMs), SHACL shape constraints, and human-i

- arXiv
- 2511.03466
- Published
- 2025-11-05
- Authors
- Ringwald Celian, Gandon Fabien, Faron Catherine, Michel Franck, Abi Akl Hanna
AI summary
Overview
Research area: Natural Language Processing — relation extraction from text into structured knowledge graphs, with a focus on small language models (SLMs), SHACL shape constraints, and human-in-the-loop active learning.
Technical level: Intermediate. Readers should be comfortable with knowledge graphs, SPARQL/SHACL, RDF triples, and the general idea of fine-tuning encoder-decoder language models, but the paper's framing is applied rather than deeply mathematical.
One-sentence scope: The paper introduces Kastor, a framework that reformulates SHACL shape-based relation extraction as selecting the best achievable property pattern per training example, then refines small fine-tuned models through a light active learning loop with a human annotator to support knowledge base completion.
What This Paper Is About
Fine-tuning small language models to extract RDF triples is attractive because it is cheap and can target a precise extraction pattern, but the classic setup forces every training example to satisfy one single maximal SHACL shape, which does not fit short texts that only partially mention an entity's properties. Kastor (Knowledge Active Shape-based extracTOR) addresses this by evaluating all combinations of properties derived from the shape and choosing the combination that actually matches each example, then adding a lightweight active learning stage in which a human corrects the model's false positives and false negatives to build a "gold" model. The goal is to produce frugal, domain-adaptable extractors whose outputs can directly populate or repair noisy knowledge bases.
Key Contributions
- A refined task definition. Kastor extends prior shape-based extraction work by reformulating validation from a single maximal SHACL shape to example-specific achievable patterns drawn from the powerset of the shape's properties, and reports better performance across a wider variety of cases.
- A generalized framework. The new task definition is embedded in an end-to-end pipeline covering sample selection, dataset construction, rule-based graph augmentation, annotation, and small-language-model fine-tuning for shape-based extractors.
- A light active learning process. A single-loop annotation workflow with a domain expert produces gold datasets and unbiased models, increasing the relevance of the produced graphs for knowledge base completion scenarios.
- An error characterization. The paper defines metrics for pattern errors (including a pattern extension capacity measure) that allow the generated triples to be checked and that could be extended to other settings. Code, models, and datasets are released openly.
Main Findings
-
Example-specific patterns increase coverage. Relaxing the requirement that every training graph validate the maximal shape
s*yields a base of example-specific patterns:|ℙ_K(s*)| = 70patterns are actually realized in the consolidated knowledge base, out of 128 possible patterns inℙ(s*), of which only 47 validates*. -
The dual base requires heavy filtering. From the 2022.09 DBpedia datadump of 6,109,994 Wikipedia abstracts and their DBpedia graphs, 1,833,493 Person graphs validate at least one of the 127 non-empty patterns. The
wikicheckfiltering step removes 40% of entities whose graph properties cannot be found in their abstract, leaving 1,093,886 entities (60%). -
Rule-based augmentation adds data cheaply. Two inference rules (
dbo:deathDateentailsdbo:deathYear;dbo:birthDateentailsdbo:birthYear), encoded as SPARQL Update queries, produced over 900,000 new triples. Per-predicate "Part Found" rates after filtering range from 35% (dbo:alias) to 94% (dbo:deathYear). -
Small training samples suffice. A 10-fold cross-validation with 900 training examples (90%), 100 test (10%) and 100 evaluation (10%) examples reached F1+ of 0.90 in 2 hours, compared with 0.97 for 4500 training examples in 13 hours. The authors report this as a good cost-performance balance.
-
Models are frugal to train. Using carbontracker, training each fine-tuned model required under 10 minutes and a small CO2 footprint: 0.018 and 7.65 minutes for
M_RD-, 0.019 and 8.11 minutes forM'_RD0, and 0.018 and 7.91 minutes forM'_RD1+. -
All models produce near-perfect syntax and correct entity URIs. Across every test set in the results table, the parseable Turtle Light rate is 0.99 and the correct subject URI rate is 1.00.
-
The baseline model fails to generalize beyond the maximal shape.
M_RD-scored highly on its own test set (RD-_test: F1- 0.994, F1+ 0.935,r_↔0.97) but dropped sharply on the corrected datasetRD2+(F1- 0.887, F1+ 0.697,r_↔0.43), revealing an apparent lack of generalization over patterns that do not strictly validates*. -
The pattern-aware model generalizes better.
M'_RD0shows good performances with an average F1 macro score of 0.91 and 0.97 at the micro level. OnRD2+it reaches F1- 0.94, F1+ 0.785 andr_↔0.72. -
The gold model is strongest on corrected data.
M'_RD1+records the best F1 metrics and low loss, with F1- 0.958 and F1+ 0.806 onRD2+, andr_↔0.95 onRD1+_test. The paper notes its rate of triples following the initial patterns is twice that ofM_RD-and 10% more thanM'_RD0. -
Active learning drives pattern extension. In the error analysis,
M'_RD1+onRD2+achieves a pattern extension capacityPEC_Dof 0.87 and produces 27.60 distinct predicted patterns against 28.90 expected, the closest match among the compared settings, which the authors describe as promising for knowledge completion. -
The annotator's corrections improve dataset quality. Corrected datasets
RD1+andRD2+contain more properties per entity (3.27 and 3.24 versus 2.8-2.9), a higher rate of graphs valid againsts*(0.59 and 0.59 versus 0.47-0.49) and higher average NLI (0.59, 0.60 versus 0.40-0.42) and Triplet Critic (0.75, 0.75 versus 0.54-0.55) scores.
Methodology in Plain English
The researchers start from a "dual base": pairs of a Wikipedia abstract and the DBpedia graph describing the same entity, taken from the 2022.09 DBpedia datadump. They pick one SHACL shape s* describing the dbo:Person class with 7 datatype properties (rdfs:label, dbo:alias, dbo:birthName, dbo:birthDate, dbo:deathDate, dbo:birthYear, dbo:deathYear). Because abstracts are often short and mention only some of these properties, the authors do not demand that every training graph satisfy the full shape. Instead, for each example they enumerate the powerset of properties derivable from the shape and keep the combination that the example's graph actually matches, calling these "example-specific patterns." They then enrich the graphs with simple inference rules (converting dates into years) so the model does not have to learn basic reasoning, and they filter out any pair where the abstract does not actually entail the graph, to reduce noise that would encourage hallucination.
The filtered base is sampled into small datasets: RD- (only graphs validating s*, used to reproduce the original baseline model M), RD0 and RD1 of 1200 examples each, and an independent control set RD2 of 600 examples. The models are BART-base (140M parameters) fine-tuned on a single Tesla V100-SXM2-32GB GPU with an inverse square root scheduler, initial learning rate 0.00005, 1000 warmup steps and early stopping with patience 5, using a prompt of the form "$entity_URI : $Abstract" and TurtleLight-linearized graphs.
For the active learning loop, a first model M'_RD0 is trained on RD0, then asked to predict graphs for RD1 and RD2. Only the false positive and false negative triples are extracted, and a domain expert judges each one against the corresponding abstract: a triple counts as erroneous if its value cannot be found in the text or does not exactly equal the expected value. Correct false positives are added and erroneous false negatives are deleted, producing gold datasets RD1+ and RD2+. RD1+ trains a gold model M'_RD1+, while RD2+ serves as an untouched control. All models are also cross-evaluated, including the gold model tested on the uncorrected RD2, and evaluated with parseability, URI correctness, loss, macro/micro F1, false positive and false negative rates, and pattern-equivalence metrics.
Why This Matters
Impact on research. The paper argues that shape-based relation extraction is a new task that cannot be directly compared with standard relation extraction benchmarks, because existing fine-tuned models offer no control over which properties are extracted and use differing linearizations, while large language models struggle with structured output despite few-shot abilities. Kastor contributes an alternative framing: extractors that obey a declared constraint, trained on limited text and RDF data, with an error taxonomy for the triples they produce. It also adds evidence to the debate about small models versus large models and about the cost of human annotation in noisy distant-supervision settings.
Real-world applications.
- Populating and repairing specialized or domain knowledge bases, where Kastor's design of coupling an incomplete, noisy base with language models is intended to uncover new relevant facts.
- Curating person-related data from text (the paper's shape targets
dbo:Person), such as biographical records with birth/death dates, years, names and aliases. - Building extraction pipelines where outputs must conform to a predefined schema, since the model emits valid, directly ingestible RDF rather than free text.
- Auditing existing knowledge bases, using the false positive / false negative triage workflow to identify facts in a graph that are unsupported by the associated text.
Industry relevance. The reported training cost is under 10 minutes per model on a single GPU with a minimal CO2 footprint, and the authors show that a 10-fold cross-validation on 900 training examples is sufficient, which lowers the barrier for organizations without large compute budgets. The human-in-the-loop step is deliberately light (one annotation loop), which matters because the paper stresses that human feedback is costly but necessary.
Future Directions
- Extending beyond the person shape. The paper reused one shape targeting
dbo:Person, which it notes represents 1/6 of the DBpedia content, and states that this shows Kastor already scales; testing the framework on other classes and other shapes is the natural next step. - Wider active learning loops. Only a single-loop annotation process is described and evaluated here; whether iterating the loop further, or combining the annotator with the LLM-as-judge and CoAnnotation ideas cited in the related work, changes the gold model's quality is left open.
- Applying the error characterization elsewhere. The paper proposes its error analysis partly because it "could be extended in other settings," implying that the pattern error and pattern extension metrics could be adapted to other extraction pipelines.
- Determining what the active process is worth end-to-end. The gold model reaches 0.87 pattern extension capacity on
RD2+, but the paper's broader conclusion, final discussion, and any explicit list of future work are not available in the truncated content provided here, so the authors' own stated next steps beyond the above are not reported.
Target Audience
Researchers and practitioners in natural language processing and knowledge graph engineering who work on relation extraction, structured output generation, or knowledge base completion. It is particularly relevant to readers interested in small, frugally trained models as alternatives to large language models, to those using SHACL shapes or RDF schemas to constrain extraction, and to teams that need to justify the cost of human annotation in their data pipelines. Readers unfamiliar with RDF, SHACL, and SPARQL will need background reading, since the framework is defined in those terms throughout.
Authors’ abstract
RDF pattern-based extraction is a compelling approach for fine-tuning small language models (SLMs) by focusing a relation extraction task on a specified SHACL shape. This technique enables the development of efficient models trained on limited text and RDF data. In this article, we introduce Kastor, a framework that advances this approach to meet the demands for completing and refining knowledge bases in specialized domains. Kastor reformulates the traditional validation task, shifting from single SHACL shape validation to evaluating all possible combinations of properties derived from the shape. By selecting the optimal combination for each training example, the framework significantly enhances model generalization and performance. Additionally, Kastor employs an iterative learning process to refine noisy knowledge bases, enabling the creation of robust models capable of uncovering new, relevant facts