Skip to content
AI.info

Research

HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon

Overview Research area: Natural Language Processing, specifically Named Entity Recognition (NER) for low-resource languages, with a focus on Hiligaynon, a Philippine regional language. Technical level

HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
arXiv
2510.10776
Published
2025-10-12
Authors
James Ald Teves, Ray Daniel Cal, Josh Magdiel Villaluz, Jean Malolos, Mico Magtira, Ramon Rodriguez, Mideth Abisado, Joseph Marvin Imperial

AI summary

Overview

  • Research area: Natural Language Processing, specifically Named Entity Recognition (NER) for low-resource languages, with a focus on Hiligaynon, a Philippine regional language.
  • Technical level: Intermediate. The methods are standard transformer fine-tuning (mBERT, XLM-RoBERTa) rather than novel architectures, but the evaluation setup, BIO tagging scheme, and cross-lingual experiments assume familiarity with NLP pipelines.
  • Scope: The paper introduces HiligayNER, described as the first publicly available NER dataset and baseline model for Hiligaynon, built from over 8,000 annotated sentences and evaluated with two fine-tuned multilingual transformer models plus zero-shot transfer to Cebuano and Tagalog.

What This Paper Is About

Hiligaynon, spoken by over 10 million people in Western Visayas (Panay Island, Negros Occidental, and Soccsksargen), has no public annotated corpus or baseline NER model, which blocks downstream information-extraction work for the language. The authors address this gap by collecting, cleaning, and manually annotating Hiligaynon text with named-entity labels, then fine-tuning two multilingual transformer models on it to establish reference performance. They also test whether models trained on Hiligaynon transfer to the related languages Cebuano and Tagalog.

Key Contributions

  1. A cleaned, sentence-level Hiligaynon dataset of over 8,000 entries (8,082 sentences after preprocessing from an initial raw collection of 17,647), sourced from publicly accessible news articles, social media posts, and translated texts across five platforms.
  2. Span-level BIO-encoded annotations of that corpus for NER, covering Person (B-PER, I-PER), Organization (B-ORG, I-ORG), Location (B-LOC, I-LOC), and Other (OTH) categories, validated with a Cohen's kappa of 0.8141.
  3. Two fine-tuned multilingual Transformer models, mBERT and XLM-RoBERTa, for token-level sequence labeling of Hiligaynon text.
  4. Release of the dataset, model checkpoints, annotation protocol, and evaluation scripts under an open license (CC BY-NC-SA 4.0), with code and data at https://github.com/jvlzloons/HiligayNER.

Main Findings

  • Annotation quality: On a stratified 10% subset annotated independently by all three annotators, observed agreement was 0.9493 and agreement by chance was 0.7273, yielding Cohen's kappa = 0.8141, which the paper interprets as substantial agreement.
  • Clean corpus size: The raw collection of 17,647 sentences was reduced to 8,082 after removing malformed strings, empty lines, and non-Hiligaynon text. Source-level counts after cleaning were 5,500 (Ang Pulong Sang Dios), 1,877 (Ilonggo News Live), 276 (Hiligaynon News and Features), 276 (Bombo Radyo Bacolod), and 153 (Ilonggo Balita sa Uma).
  • mBERT performance: Macro F1 of 0.86 at the token level. Per-tag F1 scores were 0.96 (B-PER), 0.94 (I-PER), 0.83 (B-LOC), 0.82 (I-LOC), 0.82 (B-ORG), and 0.79 (I-ORG).
  • XLM-RoBERTa performance: Per-tag F1 scores were 0.96 (B-PER), 0.94 (I-PER), 0.82 (B-LOC), 0.84 (I-LOC), 0.81 (B-ORG), and 0.79 (I-ORG).
  • Entity-type pattern: Person entities were recognized most accurately by both models (0.96 for B-PER, 0.94 for I-PER). Location entities were second best. Organization entities were the most difficult, and the paper notes that per-tag performance correlates with tag frequency, citing I-PER with 2,181 instances versus B-ORG with 505 instances.
  • Discrepancy worth noting: The abstract states that both models achieved "over 80% in precision, recall, and F1-score across entity types," but the per-tag tables report I-ORG F1 of 0.79 for both mBERT and XLM-RoBERTa.
  • Training behavior: Training loss fell monotonically over the first 100 batches and flattened afterward; validation loss stayed below 0.05, which the authors read as no over-fitting. F1 improved from 0.79 to 0.87 for mBERT and reached 0.88 for XLM-RoBERTa during training.
  • Error analysis: Confusion matrices showed that B-PER and I-PER accounted for more than 96% of their respective instances on the main diagonal, with cross-category bleed between person and non-person tags below 0.5% and false positives for rare classes below 1% of total predictions. Residual error concentrated on the ORG–LOC boundary.
  • Cross-lingual transfer: Zero-shot evaluation on Cebuano and Tagalog produced macro F1 scores around 0.46 (reported range 0.44 to 0.46). Full figures: mBERT on Cebuano (precision 0.4402, recall 0.4773, F1 0.4580, accuracy 0.9727) and Tagalog (precision 0.3998, recall 0.4991, F1 0.4439, accuracy 0.9639); XLM-RoBERTa on Cebuano (precision 0.4340, recall 0.4984, F1 0.4640, accuracy 0.9736) and Tagalog (precision 0.3894, recall 0.5221, F1 0.4461, accuracy 0.9633).
  • Evaluation scope: Only token-level precision, recall, and F1 were reported. Entity-level (span-level) evaluation was explicitly not conducted, following the convention the authors attribute to the CebuaNER study.

Methodology in Plain English

The authors worked in stages. First, they crawled five publicly available online platforms for Hiligaynon content and split the text into individual sentences, keeping only well-formed Hiligaynon material. Second, three undergraduate linguistics students who are native Hiligaynon speakers annotated the corpus in Label Studio using the CoNLL-2003 BIO convention, where a B- prefix marks the first token of an entity and an I- prefix marks subsequent tokens. Annotators received ten hours of joint training, including pilot rounds on 250 sentences adjudicated by a supervising linguist, and disagreements were settled through consensus meetings. To check reliability, all three annotators independently labeled a stratified 10% overlapping subset, and the authors computed Cohen's kappa on pairwise comparisons.

Third, they fine-tuned two pretrained multilingual encoders, mBERT and XLM-RoBERTa, using the standard token-classification pipeline in Hugging Face Transformers. A softmax classification head maps each token representation to one of the four entity tags. XLM-RoBERTa used the same recipe as mBERT but with the learning rate lowered to 3×10⁻⁵, following XLM-R recommendations. Fine-tuning ran for three epochs with AdamW. Finally, they ran zero-shot cross-lingual tests by applying the Hiligaynon-trained checkpoints to Cebuano and Tagalog data, and examined confusion matrices to see where errors landed.

Why This Matters

  • Impact on research: The work fills a documented gap — the paper states that Hiligaynon previously had no public NER corpus or baseline model, with computational work limited to tokenization heuristics and morphosyntactic lexicons. It establishes a reproducible reference point and adds a Philippine language to the set with published NER baselines, alongside prior Tagalog and Cebuano work.
  • Real-world applications:
    • Information-extraction pipelines for regional journalism, where Hiligaynon is the dominant medium in Western Visayas.
    • Public administration use cases that require mining local-language documents and announcements.
    • Social-media analytics on Hiligaynon-language posts, which the paper identifies as a major source of its data.
    • Knowledge-graph construction, information retrieval, and domain-specific analytics, which the paper cites as downstream applications of robust NER generally.
  • Industry relevance: The released checkpoints provide a starting point for adapting to related Central Philippine languages without training from scratch, and the authors note a connection to the GamotPH (General Access Multilingual Online Tool for Public Health Drug-Reporting) project, which received financial support from National University and the Department of Science and Technology.

Future Directions

  • Expand and diversify the corpus, since the current dataset was reduced from 17,647 raw sentences to 8,082 and draws from a limited set of five sources.
  • Add finer-grained entity tags, such as Event and Date, beyond the current Person, Organization, Location, and Other categories.
  • Improve organization recognition, which both models handled worst. The authors specifically recommend gazetteer augmentation and span-level objectives targeting the ORG–LOC boundary.
  • Conduct entity-level (span-level) evaluation, which the paper acknowledges is a stricter measure of system performance and explicitly leaves to future work. They also suggest domain-adaptive pretraining on regional news.

Target Audience

This paper is most useful to NLP researchers working on low-resource and underrepresented languages, particularly those focused on Philippine languages and Southeast Asian language technology. It also serves practitioners who need a ready-made Hiligaynon NER baseline or a transfer starting point for Cebuano and Tagalog, and corpus linguists or annotation teams interested in the paper's annotation protocol, native-speaker annotator training, and reported inter-annotator agreement figures.

Authors’ abstract

The language of Hiligaynon, spoken predominantly by the people of Panay Island, Negros Occidental, and Soccsksargen in the Philippines, remains underrepresented in language processing research due to the absence of annotated corpora and baseline models. This study introduces HiligayNER, the first publicly available baseline model for the task of Named Entity Recognition (NER) in Hiligaynon. The dataset used to build HiligayNER contains over 8,000 annotated sentences collected from publicly available news articles, social media posts, and literary texts. Two Transformer-based models, mBERT and XLM-RoBERTa, were fine-tuned on this collected corpus to build versions of HiligayNER. Evaluation results show strong performance, with both models achieving over 80% in precision, recall, and F1-score across entity types. Furthermore, cross-lingual evaluation with Cebuano and Tagalog demonstrates promising transferability, suggesting the broader applicability of HiligayNER for multilingual NLP in low-resource settings. This work aims to contribute to language technology development for underrepresented Philippine languages, specifically for Hiligaynon, and support future research in regional language processing.

Read the original paper