Research
PCFEx: Point Cloud Feature Extraction for Graph Neural Networks
Overview Research area: Computer vision and graph-based deep learning applied to 3D point clouds, specifically millimeter-wave (mmWave) radar data for Human Pose Estimation (HPE) and Human Activity Re

- arXiv
- 2603.08540
- Published
- 2026-03-09
- Authors
- Abdullah Al Masud, Shi Xintong, Mondher Bouazizi, Ohtsuki Tomoaki
AI summary
Overview
Research area: Computer vision and graph-based deep learning applied to 3D point clouds, specifically millimeter-wave (mmWave) radar data for Human Pose Estimation (HPE) and Human Activity Recognition (HAR).
Technical level: Advanced. The paper assumes familiarity with graph neural networks, graph attention, point cloud representations, and radar signal attributes.
Scope: The paper proposes PCFEx, a point cloud feature extraction pipeline (node, edge, and frame level) paired with a graph attention network architecture, and evaluates it on four public mmWave radar benchmarks.
What This Paper Is About
Radar point clouds contain rich per-point information, but most existing models either flatten them into image-like tensors (which distorts spatial relationships between points) or feed raw points into MLPs that treat each point independently. PCFEx treats the point cloud as a graph and explicitly computes new features describing each point, each connection between points, and the cloud as a whole, then processes those features with a graph attention network. The goal is more accurate and noise-robust HPE and HAR from mmWave radar without relying on RGB imagery.
Key Contributions
- Node feature expansion from 5D to 19D. The authors introduce node feature processing plus a
Statboxmodule (10 statistical operators: mean, standard deviation, median, skewness, kurtosis, geometric-mean, quantile at 25% and 75%, percentile at 25% and 75%) that extends each point's feature vector beyond the original x, y, z, Doppler velocity, and intensity, with scope to extend further. - Frame-level feature extraction. A parallel branch uses
Statboxto compute whole-cloud statistics, producing a 380-dimensional frame feature vector that acts like a residual shortcut path from the raw point cloud to the frame representation. - A frame-feature processing branch inside the GNN. A dedicated FCN-based block (
h_frame) accommodates frame-level features alongside the aggregated node features. - Node downsampling and broad validation. Grid-based node sampling reduces computational cost on dense clouds, and the method is validated on four public mmWave benchmarks (MARS, mRI, MMFi, MMActivity) plus the ShapeNet object point cloud dataset.
Main Findings
- HAR accuracy of 98.8%: The abstract reports an overall accuracy of 98.8% on mmWave-based HAR (MMActivity), described as outperforming existing state-of-the-art models. Table I shows 98.80 for the full Node+Edge+Frame configuration.
- Substantial HPE error reductions on mRI: The authors report 29% and 25% error reductions in MPJPE and PA-MPJPE respectively over state-of-the-art methods on the mRI dataset. S1 (random) splits showed greater error reduction than S2 (subject-wise) splits.
- Node features drive HPE gains; frame features drive HAR gains: In the ablation (Table I), adding node features (5D + 14D) dropped MMFi-S2P1 MPJPE from 75.63 mm to 71.05 mm and PA-MPJPE from 58.44 mm to 54.18 mm. For HAR, frame features raised accuracy from 95.01 to 97.72. The full Node+Edge+Frame combination was best overall: 70.63 mm / 53.83 mm on MMFi-S2P1, 68.61 mm / 54.54 mm on MMFi-S2P2, 2.454 cm MAE / 4.330 cm RMSE on MARS, and 98.80 accuracy on MMActivity.
- Cross-dataset comparison: MMFi's S2P3 MPJPE error is 72.74 mm, lower than mRI S2P1's MPJPE error of 99.83 mm. Conversely, mRI's S1 splits show lower error than MMFi's S1 splits. The environment-wise split (MMFi S3) is described as the most difficult.
- Generalization to object point clouds: On ShapeNet (55 real-world objects, 80/20 train-validation split), adding node and frame feature extraction improved classification accuracy from 74.57 to 78.18. The authors note this is below state-of-the-art results (max 85.9%) because they trained from scratch without multi-view techniques, vertex normals, surface point cloud sampling, or data augmentation.
- Compute-performance trade-off: On MMFi-S2P1, downsampling to Q=1 cut data processing time from 1h 46m to 36m, storage from 110 GB to 23.9 GB, and training time from 5h 50m to 2h 5m, at the cost of MPJPE rising from 70.9 mm to 76.83 mm and PA-MPJPE from 53.64 mm to 60.2 mm. Q=5 gave 44m, 69.6 GB, 4h 6m, 72.32 mm, and 55.93 mm respectively. Average GPU usage at batch size 32 was 3100 MB (Q=1), 13300 MB (Q=5), and 23900 MB (no downsampling).
- Model complexity: On MMFi S2P2, the authors report that their approach has fewer trainable parameters than the two benchmarks compared against but requires significantly longer feature processing and training time, while achieving a large performance improvement. The specific comparison values are not included in the provided text.
- Metric caveat: MPJPE in the paper is computed on a skeleton adjusted by the mid-hip point. The authors contacted the authors of the MMFi and mRI papers for confirmation of this adjustment and received no response; the shared MMFi code suggests MPJPE may be calculated without skeleton adjustment. PA-MPJPE is unaffected by this adjustment and is described as the more reliable metric.
Methodology in Plain English
Each radar frame's points are treated as nodes in a directed graph. Every node is connected to its K nearest neighbors by Euclidean distance; in all experiments K = 20 (illustrations in the paper use K = 8 for clarity). Consecutive frames are merged first (frame fusion) to increase point density, and dense clouds can optionally be reduced by dividing space into grid cells and randomly keeping at most Q points per occupied cell.
Three kinds of features are then computed:
- Node features (19-dimensional): the original x, y, z, Doppler velocity, and intensity, plus statistics from
Statboxand distance/direction from each point to the cloud centroid. - Edge features (6-dimensional per edge): Euclidean distance, angles along x, y, and z, relative velocity, and relative intensity between a point and its neighbor.
- Frame features (380-dimensional): column-wise statistics of the whole cloud, plus statistics of squared distances from every point to the centroid.
Node and edge features pass through shared MLP blocks (FCN layers with ReLU), then into graph attention layers that combine node and edge information and learn how much weight to give each neighbor. Dropout with ratio 0.5 is used inside the GAT layer. Node features are aggregated by average pooling (chosen after trying several options), concatenated with the processed frame features, and passed to a prediction head: either a frame-wise head or a sequential head using a bidirectional LSTM over L consecutive frames. The distance matrix computed for frame features is reused for node feature extraction, edge feature extraction, and defining the graph's edge matrix to save computation.
Training used Python, PyTorch, and PyTorch-Geometric on an Intel i7-12700K 64-bit processor, 32 GB RAM, and an Nvidia RTX-3090-ti GPU. Losses were MPJPE for mRI and MMFi, MSE for MARS, and cross-entropy for MMActivity. The MARS sequential variant used a reduced architecture with 64 units and 3 layers in all blocks except the single-layer LSTM, frame fusion F = 3, and a 16-frame input window. MMActivity used sequence length 60 frames with a 10-frame interval.
Why This Matters
Impact on research: The paper argues that the internal structure of point clouds is often ignored in prior radar-based HPE and HAR work, forcing models to learn it implicitly and making them vulnerable to noise and outliers. It positions explicit node, edge, and frame feature extraction plus graph attention as a more noise-robust and structurally aware alternative to image-like tensor representations, and demonstrates the feature extraction as a plug-and-play module applicable to CNN, Transformer, MLP, or GNN backbones.
Real-world applications:
- Privacy-preserving human monitoring in homes or care facilities, since radar avoids cameras and preserves privacy better than RGB.
- Health and rehabilitation tracking, motivated by the mRI dataset's rehabilitation activity protocol and its subject-wise evaluation setting.
- Activity recognition in smart environments, where mmWave radar is robust to lighting, body orientation, and shadows and consumes significantly less power than vision sensors.
- Deployment in resource-constrained settings, using node downsampling to trade a modest accuracy loss for large reductions in processing time, storage, and GPU memory.
Industry relevance: The complexity-versus-performance analysis gives concrete resource figures (processing time, storage, GPU memory, training time) that matter for real deployments, and the broadband benchmark coverage across four public mmWave datasets supports reproducibility.
Future Directions
- Extend to multi-person HPE and HAR, which the current model does not support (it handles a single person per frame).
- Validate on real-world mmWave data, since all datasets used were collected in controlled environments and only MMFi spans multiple environments.
- Improve generalization across environments, as the environment-wise split (MMFi S3) was the hardest to achieve low MPJPE on, and the authors state that further improvement is needed for real-world applications.
- Reduce feature processing and training cost, which the authors identify as significantly longer than the benchmarked alternatives despite fewer trainable parameters, and refine the role of frame features in sequential HPE, where they observed slight MPJPE degradation but significant HAR improvement.
Target Audience
Researchers and practitioners working on graph neural networks, 3D point cloud processing, or mmWave radar sensing for human pose estimation and activity recognition. It is also useful for engineers evaluating the resource cost of on-device radar perception pipelines, and for those interested in feature engineering strategies that transfer across point cloud domains (mmWave radar and object classification alike).
Authors’ abstract
Graph neural networks (GNNs) have gained significant attention for their effectiveness across various domains. This study focuses on applying GNN to process 3D point cloud data for human pose estimation (HPE) and human activity recognition (HAR). We propose novel point cloud feature extraction (PCFEx) techniques to capture meaningful information at the point, edge, and graph levels of the point cloud by considering point cloud as a graph. Moreover, we introduce a GNN architecture designed to efficiently process these features. Our approach is evaluated on four most popular publicly available millimeter wave radar datasets, three for HPE and one for HAR. The results show substantial improvements, with significantly reduced errors in all three HPE benchmarks, and an overall accuracy of 98.8% in mmWave-based HAR, outperforming the existing state of the art models. This work demonstrates the great potential of feature extraction incorporated with GNN modeling approach to enhance the precision of point cloud processing.