Skip to content
AI.info

Research

Improving Detection of Rare Nodes in Hierarchical Multi-Label Learning

Overview Research area: Hierarchical multi-label (HML) classification, class-imbalance handling, and uncertainty quantification in neural networks. Technical level: Advanced. The paper assumes familia

Improving Detection of Rare Nodes in Hierarchical Multi-Label Learning
arXiv
2602.08986
Published
2026-02-09
Authors
Isaac Xu, Martin Gillis, Ayushi Sharma, Benjamin Misiuk, Craig J. Brown, Thomas Trappenberg

AI summary

Overview

Research area: Hierarchical multi-label (HML) classification, class-imbalance handling, and uncertainty quantification in neural networks.

Technical level: Advanced. The paper assumes familiarity with hierarchical classification, binary cross-entropy objectives, ensemble methods, and divergence-based uncertainty measures.

Scope: The paper proposes and evaluates a weighted loss function that combines node-level imbalance weighting with an uncertainty-driven focal term, in order to improve prediction of rare (deep, fine-grained) nodes in hierarchical multi-label problems.

What This Paper Is About

In hierarchical multi-label classification, each observation can carry several labels arranged in a hierarchy, and deeper child nodes are naturally rarer than their parents. Because of this, models tend to predict only the common ancestor nodes and stop short of the fine-grained descendant nodes that are often the scientifically interesting ones. The authors' goal is to change the training objective so that the model pays more attention to rare and uncertain nodes, rather than to rare observations, and thereby reach deeper into the hierarchy.

Key Contributions

  1. Node-wise imbalance weighting. The authors examine different ways of defining and implementing imbalance weights in HML learning, arguing that weighting should be tied to node occurrence frequency rather than to observations. They provide evidence for which formulations work and which do not, and introduce a minimum-weight floor (w̃₀, described as an information gate) so common nodes still contribute useful gradients.

  2. Uncertainty-based focal weighting. The standard focal-loss confidence term is replaced with modern ensemble uncertainty quantities. The authors adapt and compare four candidate uncertainty measures as the focal term: the binary Bayesian Model Average (bBMA), Gated Margin Uncertainty (GMU), and epistemic uncertainty computed with either Kullback–Leibler or Jensen–Shannon divergence.

  3. A combined weighted objective. The two components are multiplied together and applied to the "max constraint loss" (ℒ_MC) of the C-HMCNN framework, giving a single loss that simultaneously re-weights rare nodes and emphasises uncertain nodes during training.

  4. Empirical evaluation and released code. The approach is evaluated on 16 gene product datasets (FUN and GO annotation schemes) and on a benthic imagery dataset (BenthicNet-E). Code is released at https://github.com/DalhousieAI/Hierarchical_Multi-Label_Imb_Focal_Weighted_Loss.

Main Findings

  • Recall improvements up to a factor of five. The abstract reports improvements in recall "by up to a factor of five on benchmark datasets," along with statistically significant gains in the F₁ score.

  • Node weighting with w̃₀ = 0.25 outperforms w̃₀ = 0.50 on the FUN datasets shown. In every FUN dataset row visible in Table 1, the w̃₀ = 0.25 configuration gives the highest recall (and is bolded as significantly greater than all others by more than 2σ), while w̃₀ = 0.50 consistently lands below it.

  • Concrete FUN examples (Table 1). On CELLCYCLE (FUN), recall rises from 1.74 ± 0.06 with no method to 4.59 ± 0.11 with w̃₀ = 0.25, and F₁ from 2.61 ± 0.09 to 4.85 ± 0.11. On DERISI (FUN), recall rises from 0.34 ± 0.03 to 2.16 ± 0.07 and F₁ from 0.46 ± 0.03 to 1.82 ± 0.05. On EISEN (FUN), F₁ rises from 4.82 ± 0.25 to 7.16 ± 0.16 and recall from 3.29 ± 0.20 to 6.69 ± 0.12. On SEQ (FUN), F₁ rises from 4.14 ± 0.19 to 6.60 ± 0.12 and Bin. AP from 9.95 ± 0.14 to 12.13 ± 0.15.

  • Precision and AP trade-offs. The gains in recall and F₁ frequently come alongside lower precision on some datasets — for instance, on EISEN (FUN) precision drops from 14.03 ± 0.62 to 11.27 ± 0.37 with w̃₀ = 0.25. The standard AP metric (AP) is often essentially flat or slightly lower than the unweighted baseline, which the authors attribute partly to the metric's optimism and its dataset-level rather than node-level nature.

  • Existing resampling methods provide only limited gains. Across the FUN rows shown, HROS-PD is generally close to or slightly below the unweighted baseline, while LPROS shows small improvements in F₁ but no consistent large recall gain.

  • Benefits extend to convolutional models. The abstract states that the approach also aids convolutional networks on challenging tasks, such as settings with suboptimal encoders or limited data.

  • Quantitative results for the GO datasets, for the focal-weighting comparison, and for the BenthicNet-E experiments are not included in the text available here — the paper notes that GO results appear in Appendix C and hyperparameters in Appendix A. The dataset descriptions, however, are reported (see Methodology).

Methodology in Plain English

The authors build on an existing hierarchical classification method, C-HMCNN, which enforces the rule that a parent's predicted probability must be at least as large as any of its descendants'. This is done with an adjacency matrix that encodes descendant relationships and a max operation along each row; the resulting training loss is called the max constraint loss (ℒ_MC).

On top of that loss, they multiply two weighting terms:

  1. An imbalance weight per node. Each node gets a weight based on how often it (plus its descendants) appears in the data, following the standard inverse-frequency form. Because raw inverse-frequency weights become extremely small for very common nodes, the authors rescale them between a minimum and maximum and add a floor value, w̃₀. Empirically they found that setting the number of classes equal to the total number of hierarchical nodes worked better than treating the problem as purely binary. They also found that applying weights to negatively annotated nodes did not help, so negative annotations receive a weight of 1.

  2. A focal weight based on uncertainty. Where focal loss normally down-weights examples the model is already confident about, the authors substitute ensemble-derived uncertainty. They train an ensemble of models (for imagery data, a shared encoder with multiple output heads) and derive, for each sample and node, an uncertainty value. Four measures are tested: bBMA (based on the ensemble's Bayesian Model Average, rescaled so that confident negative predictions are not mistaken for uncertainty), GMU (a signal-to-noise ratio capturing how much ensemble members agree, gated by confidence), and epistemic uncertainty computed as the average pairwise KL divergence across ensemble members, plus a Jensen–Shannon variant bounded by one. A minimum uncertainty floor (U₀) and an exponential factor k control how strongly extreme-uncertainty nodes are emphasised. The focal weights are computed without gradients, which the authors say avoids the risk of ensemble collapse.

Evaluation uses node-wise F₁, precision, and recall, plus a binarised average precision (Bin. AP, thresholded at 0.5) and the conventional AP score, which the authors criticise as overly optimistic on sparse annotations and uninformative at the node level.

The datasets are 16 gene product datasets split evenly between the FUN and GO annotation schemes. FUN has hierarchies up to depth six (root nodes counted as depth zero) with datasets typically containing 500 nodes, except Eisen (FUN) at 462; GO has over 4,000 nodes, is a directed acyclic graph rather than a tree, and spreads nodes across 14 hierarchical levels, with Eisen (GO) using 3,574 nodes. The BenthicNet-E dataset contains images of Echinodermata following the CATAMI biota tree, has 13 nodes up to depth two, and has 5,077 training and 1,650 test data points.

Why This Matters

The work targets a structural failure mode of hierarchical classifiers: they predict the broad, common categories and leave the informative fine-grained ones untouched. If rare nodes can be recalled more reliably, the outputs become usable for the expert tasks that motivated the hierarchy in the first place.

Real-world applications named in the paper:

  • Seafloor classification: detecting rare species or habitats, which can indicate important environmental changes and guide government policy on maritime economies.
  • Medical domains: identifying rare gene products through hierarchical classification to help diagnose disease.
  • Genomics and bioinformatics: annotating gene products and proteins under the FUN and Gene Ontology frameworks.
  • Benthic imagery analysis: cataloguing underwater biota and seafloor characteristics from image data.

Industry relevance: any application that relies on taxonomies — product catalogues, content moderation labels, fault taxonomies in industrial monitoring — faces the same long-tail structure. A loss-level fix that requires no change to data collection is easy to adopt, and the release of runnable code lowers the barrier further. The paper also notes that as small, hand-annotated datasets are amalgamated into larger multi-label collections, hierarchical multi-label problems are likely to become more common.

Future Directions

  • Fill in the unreported comparisons. The provided text presents the focal-weighting experiments and the GO results as separate sections (with GO results in Appendix C), and the BenthicNet-E comparison is described but its numbers are not shown here. A reader needs those tables to judge which of the four uncertainty measures (bBMA, GMU, epistemic KL, epistemic JS) is actually preferable.

  • Balance recall gains against precision losses. The clearest pattern in the FUN table is rising recall and F₁ alongside falling precision. Whether the operating point is acceptable for downstream expert use — and how the threshold or w̃₀ should be chosen per dataset — is an open tuning question.

  • Extend beyond tree and DAG benchmarks to more image tasks. The abstract claims benefits for convolutional networks under suboptimal encoders or limited data; testing this more broadly across small, locally sourced scientific image datasets is the natural next step.

  • Clarify the role of aleatoric versus epistemic uncertainty. The authors argue that only epistemic uncertainty should be worth emphasising, since aleatoric uncertainty cannot be reduced with more data. Whether the ensemble-based measures in practice isolate the epistemic component cleanly remains an open question.

Target Audience

Researchers and practitioners working on hierarchical or multi-label classification with long-tailed label distributions; machine learning engineers in scientific domains such as genomics, marine ecology, and medical diagnosis who need models that reach fine-grained leaf categories; and methodologists interested in uncertainty-weighted loss functions and ensemble-based training signals. Readers should be comfortable with neural network training objectives, ensemble methods, and precision/recall evaluation — the paper is not introductory.

Authors’ abstract

In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for more detailed or fine-grained classifications. This difficulty partly arises from the natural rarity of certain classes (or hierarchical nodes) and the hierarchical constraint that ensures child nodes are almost always less frequent than their parents. To address this, we propose a weighted loss objective for neural networks that combines node-wise imbalance weighting with focal weighting components, the latter leveraging modern quantification of ensemble uncertainties. By emphasizing rare nodes rather than rare observations (data points), and focusing on uncertain nodes for each model output distribution during training, we observe improvements in recall by up to a factor of five on benchmark datasets, along with statistically significant gains in $F_{1}$ score. We also show our approach aids convolutional networks on challenging tasks, as in situations with suboptimal encoders or limited data.

Read the original paper