Skip to content
AI.info

Research

VietJobs: A Vietnamese Job Advertisement Dataset

Overview Research area: Natural Language Processing (Vietnamese NLP, low-resource language modelling) and computational labour market analytics. Technical level: Intermediate. The paper is readable as

VietJobs: A Vietnamese Job Advertisement Dataset
arXiv
2603.05262
Published
2026-03-05
Authors
Hieu Pham Dinh, Hung Nguyen Huy, Mo El-Haj

AI summary

Overview

Research area: Natural Language Processing (Vietnamese NLP, low-resource language modelling) and computational labour market analytics.

Technical level: Intermediate. The paper is readable as a dataset paper, but interpreting the benchmark requires familiarity with LLM prompting settings (zero-shot, few-shot), parameter-efficient fine-tuning (LoRA), and standard evaluation metrics (Accuracy, Macro F1, RMSE, R²).

Scope: The paper introduces and benchmarks VietJobs, a corpus of 48,092 Vietnamese job advertisements spanning all 34 provinces and municipalities of Vietnam, used to evaluate generative large language models on job category classification and salary estimation.

What This Paper Is About

Research on recruitment language has largely focused on English and other high-resource languages, leaving Vietnamese under-resourced for computational analysis despite online recruitment platforms playing a central role in Vietnam's labour market. The authors build the first large-scale, open-access dataset of Vietnamese job advertisements to support linguistic, socio-economic, and NLP research. They then test whether current generative LLMs can reliably classify job postings into occupational categories and predict advertised salaries from structured job attributes.

Key Contributions

  1. A new public dataset: VietJobs contains 48,092 job postings and 15,429,581 tokens (words) collected from publicly accessible Vietnamese recruitment platforms in July 2025, with a vocabulary of 78,002 unique tokens and coverage of all 34 provinces and municipalities.
  2. A harmonised occupational taxonomy: Raw category labels from the source platforms were reduced from 24 distinct labels to 16 consolidated occupational domains adapted from ISCO-08, O*NET, and ESCO standards, with 71.5% of postings (34,365 of 48,092) providing explicit salary fields (salary_min, salary_max, salary_avg).
  3. Benchmarks for two core tasks: The authors evaluate 10 generative LLMs (multilingual, ASEAN-focused, and Vietnamese-specific) on job category classification (16 classes) using Accuracy and Macro F1, and on salary estimation using RMSE and R², under zero-shot, few-shot, and fine-tuned conditions.
  4. A cross-dataset comparison: Salary estimation is additionally benchmarked against the Vietnam Jobs Dataset from Kaggle (Nguyen, 2025), which contains job titles without full textual descriptions, and on a combined corpus.

Main Findings

  • Job classification is hard in zero-shot: Qwen2.5-7B-Instruct achieved the best zero-shot result with 0.31 accuracy and 0.32 Macro F1. PhoGPT-4B-Chat and Granite-3.3-8B-Instruct produced outputs that failed to match the canonical label set and scored 0 on both metrics, and SeaLLMs-v3-7B-Chat scored 0.09 accuracy and 0.07 Macro F1. Vistral-7B-Chat's classification scores are not reported (marked "-" in the results table).
  • Few-shot prompting gives the largest classification gains: Accuracy rose into the 0.4x range, led by Qwen2.5-7B-Instruct at 0.47, followed by Llama-SEA-LION-v3-8B-IT at 0.45 and Sailor2-8B-Chat at 0.44.
  • Fine-tuning did not beat few-shot for classification: Accuracy values plateaued around 0.3x across models after LoRA fine-tuning. Ministral-8B-Instruct-2410 declined slightly in both accuracy (0.11) and Macro F1 (0.04) after fine-tuning, which the authors attribute to a possible overfitting problem.
  • A single model dominates salary estimation: Llama-SEA-LION-v3-8B-IT was the most consistent performer, reaching RMSE 11.72 and R² 0.07 zero-shot on VietJobs (the only model with a positive zero-shot R² besides none other — Qwen2.5-7B-Instruct scored RMSE 14.06, R² -0.46), improving to RMSE 10.65 and R² 0.16 with few-shot prompting. When fine-tuned on both datasets it recorded RMSE 10.24 and R² 0.23 on VietJobs, and RMSE 12.40 and R² 0.16 on the combined evaluation.
  • Fine-tuning improves salary prediction: Qwen2.5-7B-Instruct went from RMSE 14.06 and R² -0.46 zero-shot on VietJobs to RMSE 11.32 and R² 0.06 after fine-tuning on both datasets.
  • The reported performance ordering for salary estimation: zero-shot < few-shot < fine-tuned on VietJobs < fine-tuned on Vietnam Jobs Dataset < fine-tuned on both datasets. The authors attribute the gains from the Vietnam Jobs Dataset and combined fine-tuning to increased data diversity and representativeness.
  • Instruction-tuned multilingual models outperformed Vietnamese-specific models: The authors conclude that large-scale multilingual pretraining remains advantageous for generalisable task understanding even in this resource-specific Vietnamese context.
  • Corpus characteristics: Average posting length is 321 tokens (mean); IT & Digital Engineering postings are the longest at 354.8 words and Languages & Translation the shortest at 275.1 words. English tokens make up 0.32% of the corpus (49,375 words). Advertised salaries range from 1 to 500M VND per month, with median minimum, average, and maximum of 10 / 13 / 15M VND. The largest category is Business, Sales & Customer Service (8,276 ads) and the smallest is Agriculture, Energy & Environment (322 ads), with Other Occupations at 345.

Methodology in Plain English

The authors crawled publicly accessible Vietnamese recruitment pages in July 2025 using the open-source Crawl4AI framework, then used GPT-4o and Gemini 2.5 with detailed task-specific prompts to extract structured fields (job title, category, salary, skills, employment conditions) from heterogeneous HTML templates. URL acquisition took roughly 4–5 hours, while full-page crawling and LLM-based extraction took over 7 days. They then normalised the source platforms' messy category labels into 16 clean occupational domains.

For the benchmarks, they tested 10 instruction-tuned models across three families: multilingual models (Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Granite-3.3-8B-Instruct, Ministral-8B-Instruct-2410), ASEAN-focused models (Llama-SEA-LION-v3-8B-IT, Sailor2-8B-Chat, SeaLLMs-v3-7B-Chat), and Vietnamese-specific models (PhoGPT-4B-Chat, BloomVN-8B-Chat, Vistral-7B-Chat). For classification, models had to output one of 16 canonical category names; anything else counted as incorrect. For salary estimation, models had to output a value in the strict "X triệu" (X million VND) format. Data were split 80/10/10 into train, development, and test sets, and fine-tuning used LoRA (rank 8, alpha 16, dropout 0.2, target modules q_proj, k_proj, v_proj, o_proj) with AdamW at a learning rate of 5×10⁻⁵, micro-batch size 4, effective batch size 64, 2 epochs, maximum sequence length 512 tokens, checkpointing every 200 steps, and BF16 precision on a single NVIDIA A40 GPU. Classification fine-tuning took about 5 hours per model; salary estimation models converged within 1–2 hours.

Why This Matters

Impact on research: VietJobs fills a documented gap in Vietnamese NLP resources and provides reusable benchmarks for two structured prediction tasks. It also complements existing recruitment corpora such as the Adzuna Global Job Listings dataset (over 17,000 English postings), the Djinni Recruitment Dataset (150,000 jobs and 230,000 anonymised candidate profiles in English and Ukrainian), and the AraJobs corpus for Arabic, extending coverage into a Southeast Asian, low-resource setting. Its negative R² values in several salary configurations provide a concrete signal that current general-purpose LLMs remain far from reliable at structured labour-market prediction in Vietnamese.

Real-world applications:

  • Labour market monitoring: tracking hiring volumes, occupational demand, and salary ranges across Vietnam's 34 provinces and municipalities.
  • Salary benchmarking tools: using the salary_min, salary_max, salary_avg fields available in 71.5% of postings to build compensation comparison services.
  • Bias and representation analysis: studying how recruitment language encodes gender, age, and appearance references, and how these correlate with wage offers.
  • Job matching and search: improving category tagging and recommendation systems for Vietnamese-language recruitment platforms.

Industry relevance: Online recruitment platforms such as TopCV are central to connecting employers and job seekers in Vietnam, and the paper shows that off-the-shelf instruction-tuned multilingual models can already reach roughly 0.4x accuracy on category classification with only a few in-context examples, without any fine-tuning. That is a practical deployment path for platforms that lack labelled data, while the fine-tuning results show the remaining accuracy headroom that domain adaptation has yet to close.

Future Directions

  • Expand coverage to additional recruitment platforms and time periods, since the current dataset comes from a single platform (TopCV) and may underrepresent some sectors or informal employment.
  • Incorporate multilingual or demographic data into the corpus.
  • Explore advanced modelling methods such as retrieval-augmented generation and domain-adaptive pretraining.
  • Conduct systematic comparisons against traditional machine learning baselines, specifically TF-IDF feature representations combined with classifiers such as Logistic Regression, to quantify the benefit of LLM-based approaches.
  • Address identified data-quality issues: inconsistent salary standardisation, rounding or omission of salary values, and duplicated or templated job description text that could affect linguistic analyses and model outcomes.

Target Audience

Researchers in Vietnamese and low-resource NLP, dataset and benchmark builders, computational social scientists studying labour markets and recruitment language, and applied machine learning engineers building job classification, salary estimation, or job-matching systems for Vietnamese-language recruitment platforms. The paper also suits economists and labour market analysts who need structured Vietnamese job posting data, and readers interested in how multilingual versus regionally specialised LLMs perform on a language-specific structured prediction task.

Authors’ abstract

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions, covering 16 occupational domains and multiple employment types (full-time, part-time, and internship). Designed to support research in natural language processing and labour market analytics, VietJobs captures substantial linguistic, regional, and socio-economic diversity. We benchmark several generative large language models (LLMs) on two core tasks: job category classification and salary estimation. Instruction-tuned models such as Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT demonstrate notable gains under few-shot and fine-tuned settings, while highlighting challenges in multilingual and Vietnamese-specific modelling for structured labour market prediction. VietJobs establishes a new benchmark for Vietnamese NLP and offers a valuable foundation for future research on recruitment language, socio-economic representation, and AI-driven labour market analysis. All code and resources are available at: https://github.com/VinNLP/VietJobs.

Read the original paper