Skip to content
AI.info

Recommender systems

Taxonomies, Knowledge Graphs, and Structured Item Relations

Integrate categories, attributes, compatibility, creators, entities, and knowledge-graph paths without turning incomplete ontologies into ground truth.

By the end you can

Example

HP puts in writing that the compatibility relation can stop being true

A compatibility edge is not a fact about two objects. It is a claim somebody maintains. It can be withdrawn while the two objects stay exactly as they were.

Printers are the clearest case, because the manufacturer put it in writing. HP uses Dynamic Security to authenticate Original HP Cartridges. Its 2020 statement said what that meant for anything else in the tray: “Further, since the dynamic security authentication process can change during the life of the printer (including in response to firmware updates), cartridges using a non-HP chip or modified or non-HP circuitry may function for a period of time, but then cease to function.”

A cartridge that prints today is not promised to print after the next update.

The arrangement eventually carried a price. On 9 December 2020 Italy's competition authority, the AGCM, closed case PS11144 with sanctions of 10 million euro against HP Inc and HP Italy S.r.l. It found that the firmware refuses to print when it recognises cartridges that are non-original “o prodotte prima di una certa data”. It found the limitations had been “rinnovate e modificate attraverso successivi aggiornamenti del firmware”, without adequate disclosure. HP had 60 days to report compliance and 120 days to change the printers' packaging.

A recommender that carried a “compatible with” edge for one of those cartridges would have been right on the day it was curated. An update it never saw made it wrong.

  • Typed relation: The edge asserted here is authentication, not physical fit. HP uses Dynamic Security to authenticate Original HP Cartridges, and a connector that mates is not the relation being tested.
  • Version dependence: HP's own words are that the authentication process can change during the life of the printer, including in response to firmware updates. The AGCM found the limitations “rinnovate e modificate attraverso successivi aggiornamenti del firmware”.
  • Incomplete knowledge: A cartridge printing today is not evidence that the edge holds tomorrow. Such cartridges, HP says, may function for a period of time, but then cease to function.
  • Provider conflict: The manufacturer's metadata and a regulator's finding disagreed. It took a national authority, closing case PS11144 on 9 December 2020, to record which claim the buyer had actually been given.
  • Decision risk: 10 million euro in sanctions against HP Inc and HP Italy S.r.l., compliance reported within 60 days, packaging changed within 120 days — because the edge looked authoritative while the operational fact behind it had moved.

Structured relations add semantics that behavior alone may not recover

Taxonomies and knowledge graphs encode categories, attributes, entities, creators, brands, compatibility, substitutes, and domain rules. They can support cold start, explanations, constraints, and relation-specific retrieval. But the structure is evidence somebody curated. It carries limits in coverage, in version, and in who vouched for it. Validate the missing edges, the categories that are too broad, and the facts a provider supplied. Do that before any of them drives exposure or a safety-critical compatibility claim.

Knowledge-graph recommenders make the relation itself the evidence. A user-item connection becomes a path through entities and relations. KPRN, published in 2019, builds path representations “by composing the semantics of both entities and relations” and leverages “the sequential dependencies within a path”. It then pools those paths to say which one carried the recommendation.

KGAT, the same year, works over the knowledge graph and the interaction graph together. Its argument: “in such a hybrid structure of KG and user-item graph, high-order relations --- which connect two items with one or multiple linked attributes --- are an essential factor for successful recommendation”. The model “recursively propagates the embeddings from a node’s neighbors (which can be users, items, or attributes)”.

The same paper reports what that structure costs and what it buys. That comparison is the one worth carrying out of this lesson. On Amazon-book the knowledge graph holds 88,572 entities, 39 relation types and 2,557,746 triplets. Those sit over 70,679 users, 24,915 items and 847,733 interactions. Last-FM holds 58,266 entities, 9 relations and 464,567 triplets. Yelp2018 holds 90,961 entities, 42 relations and 1,853,704 triplets. Against all of that, the reported improvement is single-digit: “In particular, KGAT improves over the strongest baselines w.r.t. recall@20 by 8.95%, 4.93%, and 7.18% in Amazon-book, Last-FM, and Yelp2018, respectively.”

Millions of curated triplets are what those percentages stand on. Every one of them is a claim that can be wrong, stale, or absent.

A path through the graph explains a recommendation only as well as the curation behind it, and a wrong compatibility edge becomes a confident wrong answer.

Visual

A structured-relation stack

The stack runs from what an item is down to who said so. Taxonomy, attributes, entity relations and operational relations each answer a different question. Provenance and version, at the bottom, records where every answer came from and when it stops being true.

NAICS shows what that bottom layer looks like when an institution actually runs it. Revisions are considered every five years, in calendar years ending with 2 and 7. The 2022 revision is effective for reference years beginning on or after 1 January 2022. OMB published the final 2022 decisions on 21 December 2021, in the Federal Register, updating Statistical Policy Directive No. 8. The ECPC then published a comprehensive listing of changes linking 2022 national industries back to their NAICS 2017 counterparts. A concordance is exactly what a version layer owes the systems downstream of it.

The classification is co-owned by OMB, Statistics Canada and INEGI in Mexico. Statistics Canada approved NAICS Canada 2022 Version 1.0 as a departmental standard on 30 July 2021. It calls the result “the biggest revision to NAICS since 2002”.

Re-cutting a vocabulary damages everything measured under the old one, and the issuing body says so in its own words. Commenters asked for biobased products manufacturing and renewable chemicals manufacturing to be broken out as separate industries. OMB declined to delineate them further. One reason it gave: “Further delineation would also jeopardize existing time series’ continuity.”

A recommender that joins today's categories onto last year's logs is making the change OMB refused to make. Silently, and without a concordance.

FigureLayers · 5 layers
  1. 01

    Taxonomy

    Organize items into controlled categories and hierarchical concepts.

  2. 02

    Attributes

    Represent typed values such as size, genre, ingredient, or capability.

  3. 03

    Entity relations

    Connect creators, organizations, locations, topics, and franchises.

  4. 04

    Operational relations

    Capture compatibility, availability, substitution, and constraints.

  5. 05

    Provenance and version

    Record source, confidence, validity interval, and correction process.

Analogy

A museum catalog with penciled corrections

Artist, period, material, related works: a museum catalog records all four. It supports navigation, and it keeps changing as new research and disputed attributions arrive. Knowledge graphs carry the same value and the same provisional quality.

A visitor can at least see the pencil. NAICS ships its pencil marks too, as a published listing of changes linking 2022 national industries back to their NAICS 2017 counterparts. A ranked slate shows nothing of the version, the authority, or the disputed edge that put an item on it.

Structured knowledge is governed evidence, not a finished map of the world.

Comparison

Taxonomies and knowledge graphs serve different scopes

Scope separates them. A taxonomy answers where an item belongs. A knowledge graph answers what an item is connected to. A learned relation embedding compresses both into vectors that scale well and no longer say which relation was which.

The taxonomy is the layer with a standard behind it, and the standard is unusually explicit about what its own edges do not mean. SKOS is the standard data model for publishing thesauri, taxonomies, classification schemes and subject heading systems, a W3C Recommendation since 18 August 2009. It refuses to make hierarchy compose: “Note that, to support this usage convention, the properties skos:broader and skos:narrower are not declared as transitive properties.” Applications that want the closure get separate properties, skos:broaderTransitive and skos:narrowerTransitive.

Its editors later wrote down why. “It was decided that the properties skos:broader and skos:narrower would not be transitive” — a decision made to simplify implementation, since applications need to tell parent/child links from ancestor/descendant ones. So easier to govern is not the same as free. A two-hop “belongs to category” path is not an entailment unless the publisher has said it is.

The graph column pays in maintenance for what it returns in semantics and explanation. KGAT's Amazon-book graph alone is 88,572 entities, 39 relation types and 2,557,746 triplets. On top of that structure the paper claims “interpretability benefits brought by the attention mechanism”. The embedding column keeps the scale and loses the labels. That is why it needs typed evaluation rather than the ranker's usual aggregate metrics.

FigureComparison · 3 columns

Taxonomy

Provides a mostly hierarchical vocabulary.

  • Supports navigation and category balance
  • Can force ambiguous items into rigid buckets
  • Easier to govern
  • Useful for calibration and constraints

Knowledge graph

Represents many typed entity relations.

  • Supports multi-hop retrieval
  • Handles richer semantics
  • Harder to maintain and validate
  • Useful for explanation and cold start

Learned relation embedding

Compresses structured and behavioral relations into vectors.

  • Scales scoring and retrieval
  • May blur relation semantics
  • Needs typed evaluation
  • Useful as a complementary model

Steps

Operationalize structured relations

A relation contract says what the edge asserts, who asserted it, and how long that stays true. HP's Dynamic Security statement is the shape of the field that is usually missing: the manufacturer itself says a cartridge may function for a period of time and then cease to function. An edge recorded without a validity window was never the whole claim.

Hard and soft uses are separated next, missingness is tested, and behavior is used as a check. The manufacturer's system asserted one thing and buyers observed another. The AGCM found that disagreement serious enough to close case PS11144 with sanctions of 10 million euro and a 120-day order to change the packaging.

Then provenance travels with the relation into serving. NAICS is the working model of what that costs and buys: a named directive, revisions considered every five years in calendar years ending with 2 and 7, and a published concordance back to the 2017 counterparts. Nobody has to guess which vintage a row belongs to.

FigureProcess · 5 steps
  1. 1. Define relation contracts

    Specify direction, cardinality, source, confidence, and time validity.

  2. 2. Separate hard and soft use

    Keep safety constraints distinct from ranking features.

  3. 3. Test missingness

    Evaluate coverage by item type, provider, region, and age.

  4. 4. Compare with behavior

    Inspect agreements and conflicts between graph and interaction evidence.

  5. 5. Preserve provenance

    Expose source and version in debugging and user-facing explanations.

Example

Structured-data risks

Ontology bias enters when the vocabulary is chosen. Explanation overreach enters when it is shown to a user. Neither is hypothetical, and in both cases the maintainers of the best-known structures were the ones who did the counting.

ImageNet audited the person subtree it had inherited from WordNet. That subtree holds 2,832 categories, roughly 8.3% of all ImageNet images. Twelve in-house graduate-student annotators marked 1,593 of them unsafe, meaning offensive or sensitive. “The unsafe synsets are associated with 600,040 images in ImageNet. Removing them would leave 577,244 images in the safe synsets of the person subtree of ImageNet.” ImageNet announced the removal on 17 September 2019, and the audit was published at FAT* in 2020. A borrowed vocabulary was carrying judgments its users had never inspected.

On the missing edge, the Knowledge Vault paper measured coverage on the best-known curated graph rather than warning about it. Its finding, in 2014: “71% of people in Freebase have no known place of birth, and 75% have no known nationality.” Its Table 1 puts Freebase at 40M entity instances, 35,000 relation types and 637M confident facts. Knowledge Vault's own row sits beside it: 45M entities, 4,469 relation types and 271M facts at confidence 0.9 or higher, drawn from 1.6B extracted candidate triples.

A graph that large is still mostly silence. Silence is not a negative fact.

  • Ontology bias: The classification reflects one institution's worldview or commercial incentives — ImageNet's own audit marked 1,593 of the 2,832 categories in the person subtree as unsafe, offensive or sensitive.
  • Missing-edge fallacy: Absence is interpreted as a negative fact, in a world where “71% of people in Freebase have no known place of birth, and 75% have no known nationality.”
  • Version collapse: Current relations are applied retroactively to historical training examples. That is the move OMB declined to make when it warned that further delineation “would also jeopardize existing time series’ continuity”.
  • Relation laundering: A weak provider assertion gains authority by entering the graph — a manufacturer's authentication result becomes a compatibility fact, until a regulator closes case PS11144 and says what it actually was.
  • Explanation overreach: A path is shown without confidence, source, or validity period, even where the publisher of the standard has said the hops do not compose.

Key idea

The authority gate

A structured relation may guide recommendation only to the extent justified by its source, completeness, and validity window. HP's own statement is that an authentication result can change during the life of the printer. SKOS declines to declare skos:broader and skos:narrower transitive. The Knowledge Vault paper put a number on how much of Freebase is simply unsaid. Each of those is an issuer telling you the limit of the claim, in advance. An edge whose record says none of that has told you nothing.

Give each relation only as much authority as its record can defend, and treat an edge with no recorded origin or validity window as a claim that has already lapsed.

Key takeaways