Skip to content
AI.info

Research

OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph

Overview Research area: Natural Language Processing — computational lexicography and semantic knowledge graph construction. Technical level: Intermediate. The concepts (dictionaries, WordNet, semantic

OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
arXiv
2511.18622
Published
2025-11-23
Authors
Michael J. Bommarito

AI summary

Overview

Research area: Natural Language Processing — computational lexicography and semantic knowledge graph construction. Technical level: Intermediate. The concepts (dictionaries, WordNet, semantic graphs) are broadly accessible, but the paper assumes familiarity with lexical resources and LLM-based structured generation. Scope: This paper describes OpenGloss, a synthetically generated English encyclopedic dictionary and semantic knowledge graph built from a multi-agent LLM pipeline, and reports its scale, structure, and quality characteristics.

What This Paper Is About

Building a comprehensive lexical resource — one that includes definitions, example sentences, etymology, encyclopedic context, and semantic relationships — has traditionally required either decades of expert manual curation (WordNet), stitching together many incompatible sources (BabelNet), or crowdsourcing with uneven quality control (ConceptNet). The authors ask whether modern large language models combined with strict schema validation can generate such a resource systematically, cheaply, and quickly. OpenGloss is their answer: a dataset of 537K senses across 150K lexemes, produced in under one week for under $1,000.

Key Contributions

  1. A large-scale synthetic lexical dataset. OpenGloss provides 537,000 sense definitions across 150,000 lexemes, comparable to WordNet 3.1 in vocabulary breadth while providing 4.6 times more sense definitions. Each entry integrates encyclopedic context (200–400 words, 99.7% coverage), etymological histories (97.5% coverage), usage examples (averaging 2 per sense), collocations (3–6 per part-of-speech), and semantic relationships (9.1 million edges).
  2. A reproducible generation methodology. The paper documents a four-stage multi-agent pipeline with Pydantic schema-validated LLM outputs and automated quality assurance, operating within a budget under $1,000 and 96 hours of wall-clock time. This is presented as feasible for individual research groups without institutional infrastructure.
  3. An empirical positioning against existing resources. Comparisons with WordNet, BabelNet, and ConceptNet indicate complementary rather than redundant coverage — OpenGloss and WordNet share only 38% vocabulary overlap, with each contributing distinct lexicographic priorities.
  4. A publicly released resource. The dataset is available on Hugging Face under CC-BY 4.0 at both the lexeme level (mjbommar/opengloss-dictionary) and sense level (mjbommar/opengloss-dictionary-definitions).

Main Findings

  • Scale and composition: The dataset contains 150,101 lexical entries spanning 536,829 distinct senses, averaging 3.58 senses per lexeme, with a maximum of 24 senses. It comprises 94,106 single-word entries (62.7%) and 55,995 multi-word expressions (37.3%).
  • Polysemy distribution: 14,233 lexemes (9.5%) are monosemous, 92,264 (61.5%) have 2–4 senses, 42,714 (28.5%) have 5–9 senses, and 870 (0.6%) have 10 or more senses. The median is 3 senses per lexeme. Highly polysemous lexemes include "run," "set," "make," "take," and "get."
  • Part-of-speech distribution: Nouns dominate at 278,401 senses (51.9%), followed by adjectives at 144,234 (26.9%), verbs at 90,664 (16.9%), adverbs at 12,184 (2.3%), determiners at 4,755 (0.9%), prepositions at 3,379 (0.6%), interjections at 1,908 (0.4%), pronouns at 857 (0.2%), and conjunctions at 447 (0.1%).
  • Semantic graph structure: The graph contains 9.14 million edges total. At the sense level, synonymy accounts for 1.60 million edges (30.8%), hyponymy 1.42 million (27.3%), antonymy 1.12 million (21.6%), and hypernymy 1.06 million (20.3%), totaling 5.20 million sense-level edges. At the POS level, collocations account for 3.06 million edges (77.7%) and inflections 875,673 edges (22.2%).
  • Near-universal relationship coverage: Synonym relations appear in 99.7% of senses (mean 3.0 per sense), hypernyms in 99.9% (mean 2.0), hyponyms in 98.6% (mean 2.6), antonyms in 94.0% (mean 2.1 when present), and usage examples in 99.7% (mean 2.0).
  • Degree distribution difference: The generation process produces a much less right-tailed degree distribution than WordNet, so although the graph is large, most graph algorithms run significantly faster on OpenGloss than on WordNet.
  • Quality assurance results: In a validation study of 1,000 randomly sampled entries using Claude Sonnet 4.5, 38.6% of entries were flagged for including inflected forms (e.g., "running," "dogs," "better") or proper nouns (e.g., "London," "Einstein"). The authors characterize these as successful replications of WordNet's deliberate pragmatic practices rather than quality failures. Core content quality — definitions, encyclopedic paragraphs, etymologies, and usage examples — was generally usable. Remaining flags mostly concerned semantic relationship precision, which the authors suggest reflect conservative QA standards exceeding WordNet's own practices, though genuine improvements remain possible.
  • Metadata completeness: 146,066 words (97.3%) have etymology and 149,614 words (99.7%) have encyclopedia content according to Table 1, while the body text reports 97.5% etymology coverage and 99.7% encyclopedia coverage. 150,081 lexemes (99.99%) have complete sense definitions; 20 lexemes (0.01%) are incomplete generation edge cases.
  • Comparison anchors: WordNet contains 117,000 synsets (with a figure of approximately 117,659 cited in the related work); BabelNet version 5.3 contains approximately 23 million synsets across 600 languages; ConceptNet version 5.7 contains approximately 21 million edges connecting 8 million concepts across 83 languages. OpenGloss surpasses WordNet's lexeme coverage by a factor of 1.02 with 4.59 times more sense definitions.
  • Not reported in the available content: Sections 5 and 6 of the paper (detailed analysis and quality profile discussion) are truncated after Table 3, so specific results from those sections are not available in the provided text.

Methodology in Plain English

The authors built OpenGloss in four stages using specialized LLM agents connected through the pydantic-ai framework, which enforces typed schemas so that outputs that do not fit the expected structure are rejected.

Stage 1 — Choosing the words. They started with the wamerican dictionary package (version 2020.12.07-4, containing 104,334 words), applied minimal filtering (length 3–15 characters, alphabetic only, deduplication), and retained 73,200 words (70.2% coverage). They then traversed an LLM-proposed neighbor graph seeded with everyday objects, concepts, and K-12 educational topics, adding 76,901 lexemes not in wamerican. This snowball sampling loop lets relationships discovered during generation suggest additional words. Total: 150,101 lexemes.

Stage 2 — Generating senses. A two-agent design: an overview agent determines valid parts of speech, stopword classification, and approximate sense counts; a POS details agent then generates 1–4 sense definitions per part of speech, each with semantic relationships (3–5 synonyms, antonyms when applicable, 2–4 hypernyms/hyponyms, 1–3 usage examples), morphology, and 3–6 collocations.

Stage 3 — Building the graph. This stage uses no LLM calls. Edges are extracted deterministically from the structured sense data into 13 relationship types across three categories: sense-level (synonym, antonym, hypernym, hyponym), lexeme-level (collocations and morphological derivations), and historical (cognates, morpheme components, etymology parents). Edges are also given priority classifications (high/medium/low).

Stage 4 — Enrichment. Two more agents add etymological histories for non-stopword lexemes (97.5% coverage) and 200–400 word encyclopedia entries (99.7% coverage).

Infrastructure and validation. Generation used OpenAI's gpt-5-nano; quality assurance used Anthropic's Claude Sonnet 4.5. Sampling parameters were temperature 0.7 for overview, POS, and etymology and 0.9 for encyclopedia, top-p 0.95, max tokens 2048 for overview/POS and 4096 for encyclopedia, and a frequency penalty of 0.3 (partly to reduce verbatim reproduction from training data). Malformed outputs — estimated at 2–4% of responses — triggered automatic retry. Semantic validation enforced 100% edge target validity, and automated checks ensured acyclic hypernym/hyponym relationships and symmetric synonym/antonym pairs. The pipeline is backend-agnostic, so different models or deployment environments can be substituted.

Why This Matters

Research impact: The paper argues that systematic structured generation complements rather than replaces existing approaches — manual curation for gold-standard quality, integration for multilingual coverage, crowdsourcing for commonsense knowledge, and now systematic generation for comprehensive lexical resources with rapid update cycles. The dual-perspective QA framework (evaluating both against traditional lexicographic standards and against pragmatic design goals) is offered as a generalizable validation method for other synthetic structured-data projects.

Real-world applications:

  • K-12 education and vocabulary learning, where definitions, examples, collocations, encyclopedias, and etymology are integrated in one place rather than scattered across resources.
  • Reading comprehension support, with proper nouns and inflected forms included so users can look up words as they encounter them.
  • Natural language processing tasks such as word sense disambiguation, semantic similarity, and controllable text generation.
  • Graph-based reasoning applications that traverse hypernym chains or hyponym edges, potentially running faster than on WordNet given the flatter degree distribution.

Industry relevance: The under-$1,000, 96-hour cost and time profile makes comprehensive lexical resources accessible to individual research groups and small teams without institutional infrastructure, and the backend-agnostic pipeline allows rapid regeneration as foundation models improve.

Future Directions

  • Re-generation as foundation models improve. The pipeline is explicitly designed for rapid iteration; the authors present OpenGloss v1.0 as a snapshot tied to gpt-5-nano and Claude Sonnet 4.5, and note that substituting other models or open-weight alternatives is straightforward.
  • Refining semantic relationship precision. The QA study identified borderline hypernyms and synonyms as the primary remaining flag category, with the authors noting that genuine improvements remain possible there.
  • Extension beyond the pedagogical English scope. Lexeme selection was deliberately bounded to K-12 and everyday vocabulary rather than comprehensive English coverage, leaving broader coverage open, as are the multilingual directions surveyed in the related work.
  • Determining appropriate use cases. The paper states that Section 6 discusses quality profiles, validation results, and appropriate use cases in detail, but that discussion is not available in the provided content.

Target Audience

This paper is for NLP researchers and practitioners working on lexical resources, semantic knowledge graphs, and LLM-based structured generation; computational linguists interested in how synthetic resources compare to WordNet, BabelNet, and ConceptNet; and educational technology developers seeking an openly licensed, integrated dictionary resource. The paper is also relevant to researchers interested in schema-validated multi-agent pipelines as a general methodology for large-scale synthetic data creation.

Authors’ abstract

We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. OpenGloss contains 537K senses across 150K lexemes, on par with WordNet 3.1 and Open English WordNet, while providing more than four times as many sense definitions. These lexemes include 9.1M semantic edges, 1M usage examples, 3M collocations, and 60M words of encyclopedic content. Generated through a multi-agent procedural generation pipeline with schema-validated LLM outputs and automated quality assurance, the entire resource was produced in under one week for under $1,000. This demonstrates that structured generation can create comprehensive lexical resources at cost and time scales impractical for manual curation, enabling rapid iteration as foundation models improve. The resource addresses gaps in pedagogical applications by providing integrated content -- definitions, examples, collocations, encyclopedias, etymology -- that supports both vocabulary learning and natural language processing tasks. As a synthetically generated resource, OpenGloss reflects both the capabilities and limitations of current foundation models. The dataset is publicly available on Hugging Face under CC-BY 4.0, enabling researchers and educators to build upon and adapt this resource.

Read the original paper