Skip to content
AI.info

Research

DK-Root: A Joint Data-and-Knowledge-Driven Framework for Root Cause Analysis of QoE Degradations in Mobile Networks

Overview Research area: Machine learning for mobile network operations — specifically root-cause analysis of Quality of Experience (QoE) degradations from multivariate network KPI time series. Technic

DK-Root: A Joint Data-and-Knowledge-Driven Framework for Root Cause Analysis of QoE Degradations in Mobile Networks
arXiv
2511.11737
Published
2025-11-13
Authors
Qizhe Li, Haolong Chen, Jiansheng Li, Shuqi Chai, Xuan Li, Yuzhou Hou, Xinhua Shao, Fangfang Li, Kaifeng Han, Guangxu Zhu

AI summary

Overview

Research area: Machine learning for mobile network operations — specifically root-cause analysis of Quality of Experience (QoE) degradations from multivariate network KPI time series.

Technical level: Advanced. The paper assumes familiarity with supervised contrastive learning, denoising diffusion probabilistic models, and cellular protocol-layer terminology (PHY, MAC, RLC, PDCP; RSRP, SINR).

Scope: A three-stage framework (DK-Root) that combines large-scale noisy rule-based labels with a small set of expert-verified labels to classify the root causes of high-latency QoE degradations in a real operator dataset.

What This Paper Is About

Mobile operators need to know why a user's experience degrades — interference, weak coverage, or channel overload — but the KPI traces that reveal this are complex, cross-layer, and hard for humans to label. Rule-based heuristics can produce labels at scale but are noisy and coarse, while expert annotations are accurate but scarce and expensive. DK-Root's goal is to use both supervision sources together: learn broad representations from abundant noisy labels, then sharpen the decision boundary with the few trusted expert labels.

Key Contributions

  1. A joint data-and-knowledge framework. DK-Root explicitly decouples feature representation learning (driven by scalable rule-based labels) from final classification (refined by scarce expert-verified labels), so the model benefits from both data abundance and human expertise.

  2. Conditional diffusion-based generative augmentation. A class-conditional diffusion model synthesizes KPI sequences that preserve root-cause semantics. Expert labels are embedded and fused with KPI representations through a cross-modal fusion network to form semantically aligned conditioning inputs. Controlling the reverse diffusion timestep produces weak augmentation (small timesteps, stability) and strong augmentation (large timesteps, diversity).

  3. Data-aware contrastive pretraining with knowledge-based refinement. A supervised contrastive objective pulls together samples sharing the same rule-based label and pushes apart different root causes, denoising the noisy supervision. This is followed by expert-guided fine-tuning of a lightweight classification head to refine decision boundaries.

  4. Validation on a real, operator-grade dataset. Evaluation covers comparison against machine learning and deep learning baselines, qualitative inspection of generated samples, and ablation studies.

Main Findings

  • Superior accuracy claimed, but no numeric results appear in the available text. The abstract states DK-Root achieves "state-of-the-art accuracy" and surpasses traditional ML and recent semi-supervised time-series methods, and that ablations confirm the necessity of conditional diffusion augmentation and the pretrain–finetune design. The truncated content cuts off in the middle of the baseline description, so specific accuracy values, per-baseline comparisons, and ablation numbers are not reported in the material available here.

  • Dataset scale: 3,000 samples with rule-based labels (generated automatically by structured judging criteria on KPIs) and 185 samples with manually annotated expert labels. The data was collected by Huawei from real-world communication sensors at 5-second intervals, with a fixed sequence length of 40 — that is, 200 seconds per sample.

  • Task definition: A six-class classification problem over root causes of high latency: Uplink Interference, Uplink Weak Coverage, Downlink Interference, Downlink Weak Coverage, Traffic Channel Overload, and Control Channel Overload.

  • Feature structure: Each sample is a matrix $\mathbf{X} \in \mathbb{R}^{m \times l}$, where $m$ is the number of KPIs aggregated from the PHY, MAC, RLC, and PDCP layers and $l$ is the temporal length of the monitoring window.

  • Motivation for generative augmentation: The paper argues that generic time-series augmentations (jittering, scaling, time-warping) are task-agnostic and risk violating physical or protocol constraints — e.g., injecting noise into RSRP or shuffling SINR sequences can disrupt signal continuity or scheduling semantics, causing semantic drift.

  • Evaluation metric: Classification accuracy, defined as the proportion of correctly predicted samples.

Methodology in Plain English

Stage I — Training a conditional diffusion model. The researchers train a denoising diffusion model (DDPM-style) on the expert-labeled samples only. Because expert labels are noise-free, they act as an anchor for the true class manifold. The class label is turned into a continuous embedding, broadcast across the time axis, concatenated with the noisy input along the channel dimension, and fused with a 1-D convolution before entering a U-Net backbone that predicts the noise. The training loss minimizes the mean squared error between predicted and true injected noise. This conditioning lets the model learn class-specific distributions rather than one undifferentiated distribution.

Stage II — Generating views and contrastive pretraining. The diffusion model is frozen and switched to inference. For each rule-based labeled sample, two diffusion timesteps are sampled from non-overlapping subranges: $t_{\text{weak}} \sim \mathcal{U}(1, \alpha T)$ and $t_{\text{strong}} \sim \mathcal{U}(\beta T, T)$, with $0 < \alpha \leq \beta < 1$. Rather than generating from pure noise, the process starts from a real KPI sample, injects noise at the sampled timestep, and performs a single-step reverse update to produce an augmented view. The two views pass through a shared encoder, are flattened and $\ell_2$-normalized, and are trained with a supervised contrastive loss where any two samples with the same rule-based label form a positive pair. This is what denoises the noisy supervision — the model pulls same-class samples together without directly trusting the noisy labels via cross-entropy.

Stage III — Fine-tuning. A lightweight classification head is attached to the encoder from Stage II, and the whole model is fine-tuned end-to-end on the 185 expert-labeled samples using standard multi-class cross-entropy with a softmax over six classes. This stage adapts the broad representations to high-precision decision boundaries.

Why This Matters

Impact on research. The paper targets a well-known bottleneck in applied machine learning for networks: labels are either abundant-but-wrong or right-but-rare. It proposes treating the two supervision sources as complementary signals with different roles rather than mixing them into a single loss. It also makes a domain-specific argument against generic time-series augmentation, replacing it with a generative model conditioned on the fault class — a pattern that could transfer to other domains where signal semantics are physically constrained.

Real-world applications:

  • Automated network fault triage. Operators could route degradation alerts to the correct remediation team (RF optimization for coverage, interference mitigation, capacity expansion for overload) without manual KPI inspection.
  • Self-healing and autonomous operation. The paper frames this as feeding the 6G vision of autonomous operation and embedded intelligence, and references statistical digital twins for closed-loop integration between physical infrastructure and virtual models.
  • Reducing expert labeling cost. If noisy rule-based labels can be denoised via contrastive learning, operators can scale diagnostics with far fewer expensive expert annotations.
  • Cross-layer diagnostics. Because samples span PHY, MAC, RLC, and PDCP layers, the framework can surface degradation signatures that span protocol boundaries, which threshold rules handle poorly.

Industry relevance. The dataset comes from Huawei with co-authors from Huawei's Experience Lab and China Telecommunications Group, and the work is situated explicitly in the 6G autonomous-network context — a strong signal that the problem statement reflects real operational constraints rather than a benchmark abstraction.

Future Directions

  • Beyond high-latency scenarios. The paper scopes its experiments to high-latency QoE degradation; whether the framework generalizes to low-throughput, drop-rate, or mobility-induced degradations is not reported.
  • Scaling the expert-labeled set. With only 185 expert samples, the sensitivity of Stage III fine-tuning to the size and composition of the expert subset is an open question the truncated content does not answer.
  • Hyperparameter selection for augmentation strength. $\alpha$ and $\beta$ govern the weak/strong split and are described only qualitatively ("$\alpha$ is set to a small value," "$\beta$ is set to a larger value"). How sensitive results are to these choices, and whether they can be tuned automatically, is not addressed in the available text.
  • Extending the conditional fusion design. The paper notes the label embedding uses early fusion; whether richer cross-modal fusion between expert-label embeddings and KPI representations would improve generation quality is left open.

Target Audience

Researchers and practitioners working at the intersection of machine learning and telecommunications — particularly those doing semi-supervised time-series classification, generative data augmentation, or automated network operations (AIOps for RAN). It is also relevant to engineers at operators and equipment vendors building diagnostic tooling who need to understand when rule-based labeling is or is not sufficient. Readers without a background in contrastive learning or diffusion models will need to consult the cited references (DDPM, supervised contrastive learning, TS-TCC/TS2Vec) to follow the methodology in depth.

Authors’ abstract

Diagnosing the root causes of Quality of Experience (QoE) degradations in operational mobile networks is challenging due to complex cross-layer interactions among kernel performance indicators (KPIs) and the scarcity of reliable expert annotations. Although rule-based heuristics can generate labels at scale, they are noisy and coarse-grained, limiting the accuracy of purely data-driven approaches. To address this, we propose DK-Root, a joint data-and-knowledge-driven framework that unifies scalable weak supervision with precise expert guidance for robust root-cause analysis. DK-Root first pretrains an encoder via contrastive representation learning using abundant rule-based labels while explicitly denoising their noise through a supervised contrastive objective. To supply task-faithful data augmentation, we introduce a class-conditional diffusion model that generates KPIs sequences preserving root-cause semantics, and by controlling reverse diffusion steps, it produces weak and strong augmentations that improve intra-class compactness and inter-class separability. Finally, the encoder and the lightweight classifier are jointly fine-tuned with scarce expert-verified labels to sharpen decision boundaries. Extensive experiments on a real-world, operator-grade dataset demonstrate state-of-the-art accuracy, with DK-Root surpassing traditional ML and recent semi-supervised time-series methods. Ablations confirm the necessity of the conditional diffusion augmentation and the pretrain-finetune design, validating both representation quality and classification gains.

Read the original paper