Skip to content
AI.info

Research

Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies

Overview Research area: Natural Language Processing, specifically entity extraction and entity linking of job vacancy text to European labor-market and education taxonomies (ESCO and EQF). Technical l

Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies
arXiv
2512.03195
Published
2025-12-02
Authors
Stylianos Saroglou, Konstantinos Diamantaras, Francesco Preta, Marina Delianidi, Apostolos Benisis, Christian Johannes Meyer

AI summary

Overview

Research area: Natural Language Processing, specifically entity extraction and entity linking of job vacancy text to European labor-market and education taxonomies (ESCO and EQF).

Technical level: Intermediate. The paper is written accessibly, but readers will benefit from familiarity with sentence embeddings, transformer encoders, BIO-tag sequence labeling, and retrieval-based evaluation metrics like Accuracy@1.

Scope in one sentence: The paper benchmarks two approaches — sentence linking and entity linking — for mapping job vacancy text to ESCO occupations, ESCO skills, and EQF qualification levels, releases companion datasets and an open-source tool, and reports state-of-the-art entity extraction on the Green Benchmark.

What This Paper Is About

Most job-domain NLP work focuses narrowly on extracting "skills" from job postings, leaving occupations and qualifications largely unaddressed. The paper asks how language models can align the language of job advertisements with institutional classifications of labor and education — using the ESCO taxonomy (version 1.1.1, with 3,007 Occupations and 13,896 Skills) and the European Qualifications Framework (EQF, 814 English-language entries) as the target knowledge bases. The goal is both to compare competing architectures and to supply datasets, evaluation infrastructure, and an open-source classifier for future labor-market research.

Key Contributions

  1. A comparative evaluation of two methodologies for linking job vacancy text to ESCO and EQF: (Methodology 1) sentence linking, framed as extreme multi-label classification over whole job descriptions, and (Methodology 2) entity linking, which adds an intermediate entity recognition step that detects mention spans before disambiguating them against the knowledge base.
  2. Three new datasets (described as "three novel datasets" in the introduction and "two novel datasets" in the conclusion): one for occupation linking to ESCO, one for qualification linking to EQF, and one supporting occupation title similarity (210,175 title pairs covering 1,156 ESCO occupations).
  3. An open-source tool and codebase implementing both methodologies, released at https://github.com/tabiya-tech/tabiya-livelihoods-classifier, to lower the entry barrier for labor classification research.
  4. State-of-the-art entity extraction on the Green Benchmark dataset, with a RoBERTa base model reaching a strict F1-score of 54.3 ± 2.6 — surpassing the previously reported best of 51.2 by Zhang et al. (2023) — plus an investigation of whether generative large language models (GPT-4, Gemini 1.5 Pro, Universal-NER) help with this task.

Main Findings

  • Sentence linking wins for occupations. The best sentence-linking configuration (Single embedding: concatenation of preferred label and description) reached 0.4981 Accuracy@1, outperforming entity linking (0.4704) at the sentence level. The authors attribute this to SL's ability to contextualize information from the full job description.
  • Title linking is the strongest occupation strategy overall. Using only job titles as queries produced a 0.5387 Accuracy@1, an approximate 4% improvement over full-text queries, and is described as the most effective strategy when the job title is available within the job description.
  • Entity linking wins decisively for skills. EL reached 0.3969 at the sentence level versus 0.2211 for sentence linking. Adding broader context was found to introduce noise; linking based exclusively on the relevant entity text worked best. The best entity-level skill configuration used roberta-base+CRF with all-MiniLM-L6-v2 (0.326).
  • No clear winner for qualifications. EL and SL both produced 0.2881 in the sentence-level comparison, and the qualitative analysis revealed no clear advantage for either method. The authors suggest a supervised classification approach may be most appropriate, and note the terminology used for qualifications in the UK job market as a possible cause.
  • Best entity-level results per type: Occupations 0.489 (roberta-large+CRF with fine-tuned all-mpnet-base-v2), Skills 0.326 (roberta-base+CRF with all-MiniLM-L6-v2), EQF 0.350 (bert-large-cased with all-MiniLM-L6-v2).
  • Fine-tuning cuts both ways. Fine-tuning the sentence transformer on occupation-specific title data significantly improved the Occupations task without hurting Qualifications, but caused a performance drop for Skills, which the authors describe as suggestive of catastrophic forgetting.
  • CRF decoding helps unevenly. Adding a conditional random field decoder improved both BERT and DeBERTa models but did not improve RoBERTa, which remained the best-performing entity recognizer on the Green Benchmark.
  • Generative LLMs underperform supervised extractors. Gemini 1.5 Pro achieved an F1 of 0.22 with one-shot prompting and 0.25 with five-shots; Universal-NER reached 0.33 — both severely below supervised methods. Directly linking job descriptions to ESCO with LLMs was concluded to be impossible at the time due to hallucinated ESCO codes and labels.
  • Synthetic query generation did not help. Prompting Gemini 1.5 Pro to generate natural user-style queries from occupation and skill descriptions produced no improvements in retrieval results.
  • Qualification annotation is hard. Inter-annotator agreement on EQF levels measured 0.45 using Cohen's Kappa (moderate agreement); 361 of the qualification entities were labeled UNK in the evaluation set, out of 595 entities across 448 data points.

Methodology in Plain English

The researchers treat the task as a retrieval problem. Every ESCO occupation, ESCO skill, and EQF qualification is converted into a numerical vector (an embedding) using a sentence transformer, and these vectors are cached in separate vector databases. When new job text arrives, it is embedded the same way and the system returns the closest reference entries by cosine similarity, evaluated with Accuracy@1.

For sentence linking, the entire job description — or just the job title, in the title-linking variant — is used directly as the query. The team tested five ways of embedding the ESCO nodes: a single embedding from the preferred label, from the description, or from their concatenation; or multiple embeddings, one per field, either grouping secondary labels or splitting them individually. To improve occupation matching, they fine-tuned the all-mpnet-base-v2 sentence transformer on a title similarity dataset using Multiple Negatives Loss.

For entity linking, the pipeline is split in two. First, a BERT-family encoder performs token classification, assigning BIO labels to detect entity spans (this is the entity recognition module, trained on the Green Benchmark). Post-processing cleans up common sequence errors — removing special tokens like [SEP] and [CLS], fixing malformed BIO transitions, and dropping stray "I-" tags at sentence ends. Second, each extracted mention is embedded and matched against the reference sets, but only within the entity category the recognizer predicted. The extracted entities act as the queries.

Because sentence-level and entity-level evaluations are not directly comparable, the authors aggregate EL's per-entity predictions into a sentence-level list so the two methods can be compared fairly. They note that the skills evaluation set contains 920 sentence-level queries versus 2406 entity-level queries.

Why This Matters

Impact on research: The paper broadens the job-domain NLP agenda beyond skill extraction, provides two new human-annotated evaluation sets for occupations and qualifications, releases a title similarity dataset, and establishes a new state-of-the-art entity extraction result on a widely used benchmark. It also makes a candid, useful negative result: generative LLMs did not improve this task. The authors explicitly flag the relatively low annotation agreement on qualifications as a limitation of their own dataset and call for research on how to overcome it.

Real-world applications:

  • Job matching platforms that automatically align vacancy postings with standardized occupation codes.
  • Public employment services and labor market analytics that track demand for occupations, skills, and qualification levels across regions.
  • Curriculum and vocational training design, by comparing the qualifications employers request against those offered in national frameworks.
  • Retrieval-augmented generation systems in career advising, where the linking step serves as the retrieval component of a larger assistant.

Industry relevance: The work targets practitioners building labor-market products. It studies a real-world Ethiopian job dataset alongside the European frameworks, uses only open-source models except Gemini, and ships a ready-to-use codebase — a deliberate choice of transparency, reproducibility, and accessibility over maximum achievable performance.

Future Directions

  • Hybrid and constrained retrieval. The conclusion suggests future systems may benefit from combining both methods rather than choosing one.
  • Expanding qualification data. The authors recommend extending the evaluation set to qualifications from more diverse and general job markets, since the current evaluation reflects UK terminology.
  • Multilingual and non-European job markets. The work is primarily on English datasets, and ESCO is designed for Europe, which may not precisely capture low- and middle-income country job markets where occupations and idioms may not exist in the taxonomy.
  • Closing the AIDA-style data gap. The lack of a comprehensive, jointly annotated entity linking dataset for job descriptions to taxonomies is identified as a significant limitation that hinders joint training across diverse job domains.
  • Prompt transformation for retrieval. The authors suggest there may exist a prompt transformation that enhances information retrieval with synthetic queries, calling this a promising direction.

Target Audience

Researchers and practitioners in NLP applied to labor markets — particularly those working on job posting classification, skill/occupation extraction, entity linking to taxonomies, and retrieval-augmented generation. It is also relevant to labor economists, employment agencies, and policy analysts who need to understand the technical trade-offs behind automated job classification, and to annotation teams who will recognize the qualification-level disagreements the paper documents.

Authors’ abstract

This study investigates the potential of language models to improve the classification of labor market information by linking job vacancy texts to two major European frameworks: the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy and the European Qualifications Framework (EQF). We examine and compare two prominent methodologies from the literature: Sentence Linking and Entity Linking. In support of ongoing research, we release an open-source tool, incorporating these two methodologies, designed to facilitate further work on labor classification and employment discourse. To move beyond surface-level skill extraction, we introduce two annotated datasets specifically aimed at evaluating how occupations and qualifications are represented within job vacancy texts. Additionally, we examine different ways to utilize generative large language models for this task. Our findings contribute to advancing the state of the art in job entity extraction and offer computational infrastructure for examining work, skills, and labor market narratives in a digitally mediated economy. Our code is made publicly available: https://github.com/tabiya-tech/tabiya-livelihoods-classifier

Read the original paper