Research
Edge-aware baselines for ogbn-proteins in PyTorch Geometric: species-wise normalization, post-hoc calibration, and cost-accuracy trade-offs
Edge-aware baselines for ogbn-proteins in PyTorch Geometric Overview Research area: Graph machine learning — specifically graph neural networks (GNNs) applied to the Open Graph Benchmark (OGB) ogbn-pr

- arXiv
- 2511.13250
- Published
- 2025-11-17
- Authors
- Aleksandar Stanković, Dejan Lisica
AI summary
Edge-aware baselines for ogbn-proteins in PyTorch GeometricOverview
Research area: Graph machine learning — specifically graph neural networks (GNNs) applied to the Open Graph Benchmark (OGB) ogbn-proteins protein function prediction task.
Technical level: Intermediate. Readers should be comfortable with message passing, normalization layers, and multi-label classification metrics; the paper itself introduces no new architecture and works entirely with standard PyTorch Geometric components.
Scope: A one-sentence scope: this paper systematically measures how three practical design choices — aggregating 8-D edge evidence into node inputs, weighting messages by scalarized edge channels, and choosing between LayerNorm/BatchNorm/Conditional LayerNorm — plus post-hoc calibration, shift the accuracy-versus-cost frontier for GraphSAGE and GIN on ogbn-proteins.
What This Paper Is About
On ogbn-proteins, nodes are proteins, edges carry an 8-D vector of association evidence, and each protein has 112 binary functional labels that must be predicted. The official evaluation is mean ROC-AUC across labels under a species-wise split, meaning validation and test contain species that were never seen during training. Rather than proposing a new architecture, the authors ask two practical questions: how should the 8-D edge evidence be turned into node inputs, and how should edges be used inside message passing — and they report the compute cost (time, VRAM, parameters) alongside accuracy and decision quality for each choice.
Key Contributions
-
Reproducible edge-aware baselines for
ogbn-proteinsin PyTorch Geometric. The strongest baseline is GraphSAGE withsum-based edge-to-node features, trained under a standardized artifact format (args.json,metrics.json, andlogits_{val,test}.npzcontainingnode_id,species_id,logits[112],labels[112]), with public scripts at https://github.com/SV25-22/ECHO-Proteins. -
A controlled comparison of three system choices. The paper isolates (i)
meanvs.sumvs.maxaggregation of incident 8-D edge features, (ii)sumvs. a learned 1-D edge scalarizer for weighting messages in SAGE/GIN, and (iii) LayerNorm (LN) vs. BatchNorm (BN) vs. a species-aware Conditional LayerNorm (CLN) whose affine parameters are predicted from a per-species descriptor (cln_mode=desc). -
Post-hoc calibration paired with per-label thresholds. Per-label temperature scaling (plus a single global-temperature variant) is fit on validation logits by minimizing negative log-likelihood, and per-label thresholds are derived by maximizing F-beta on validation probabilities with a ROC-based fallback. Expected Calibration Error (ECE) and Brier score are reported alongside ROC-AUC and F1 — metrics that most prior
ogbn-proteinsreports omit. -
A leakage-free label-correlation smoother. A label–label co-occurrence matrix P over K = 112 labels is computed from training labels only, rows are normalized to sum to one, and logits are adjusted as z' = z + λzPᵀ with a small λ tuned on validation (λ = 0.1 in the reported runs).
Main Findings
-
sumbeatsmeanandmaxfor edge-to-node features. In the ablation with SAGE+LN (h=512, L=3),sumgives validation AUC 0.855 / test AUC 0.789 / test F1 0.144;meangives 0.839 / 0.775 / 0.138;maxgives 0.824 ± 0.000 / 0.779 / 0.141. In the MLP baselines, SUM aggregation dominates the top four configurations, reaching 0.74+ test AUC versus 0.69–0.72 for others. -
BatchNorm attains the best ROC-AUC; CLN matches the AUC frontier with better fixed-threshold F1. With SAGE (
sumedge-to-node, 3 layers, hidden 512,sumedge scalar, 3 seeds): SAGE+BN reaches val AUC 0.864 / test AUC 0.792 with test F1 0.096 and 0.650 M parameters; SAGE+CLN reaches val 0.855 / test 0.792 with test F1 0.145 and 1.710 M parameters; SAGE+LN reaches val 0.855 / test 0.789 with test F1 0.144 and 0.650 M parameters. -
A learned 1-D edge scalarizer does not beat simple
sum. SAGE+LN with a learned 1-D scalarizer gives val AUC 0.850 / test AUC 0.787 / test F1 0.128, versus val 0.855 / test 0.789 / F1 0.144 for thesumscalarizer. -
GIN lags SAGE on AUC but has a higher fixed-threshold F1. GIN+BN (sum, 3 layers, hidden 512,
sumscalar) reaches val AUC 0.845 / test AUC 0.761 with test F1 0.149 and 0.852 M parameters — lower AUC than SAGE's roughly 0.79 test AUC, but a higher fixed-threshold F1 than SAGE+BN (0.149 vs. 0.096), which the authors attribute to threshold/calibration sensitivity. -
Message passing still adds value over edge aggregation alone. Across 33 MLP configurations varying aggregation, normalization (BatchNorm, LayerNorm, none) and depth (uniform, tapering, deep), the best MLP reaches 0.743 test AUC (sum_deep_none, 847 MB VRAM, 0.181 M parameters), showing that edge aggregation alone captures significant signal but falls short of the SAGE baselines. Other MLP results include sum_taper_bn (0.743 test AUC, 0.128 F1, 1075 MB, 0.185 M), sum_deep_bn (0.742, 0.125, 1073 MB, 0.183 M), sum_uniform_256_bn (0.740, 0.137, 732 MB, 0.098 M), max_uniform_256_cln_desc (0.715, 0.120, 1021 MB, 0.098 M), mean_uniform_256_none (0.577, 0.075, 603 MB, 0.097 M) and the worst case sum_taper_none (0.465, 0.049, 848 MB, 0.183 M).
-
Post-hoc per-label temperature scaling plus per-label thresholds substantially improves micro-F1 at essentially unchanged AUC. For example, SAGE+LN moves from 0.14 micro-F1 at a fixed 0.5 threshold to roughly 0.80, with AUC changing only negligibly. In the calibration table, SAGE+LN (sum) reaches AUC 0.792 / F1 0.795 / ECE 0.188, SAGE+CLN(desc) (sum) reaches 0.795 / 0.786 / 0.183, and SAGE+BN (sum) reaches 0.795 / 0.592 / 0.350.
-
Label-correlation smoothing yields small, consistent additional gains. With λ = 0.1 in logit space (conditional-centered graph): SAGE+LN 0.794 AUC / 0.796 F1 / 0.178 ECE; SAGE+CLN 0.796 / 0.787 / 0.173; SAGE+BN 0.794 / 0.587 / 0.348.
-
BN is the least calibrated family even after post-hoc correction, while SAGE+CLN and SAGE+LN benefit most from smoothing (better ECE). In the MLP study, SUM + BatchNorm models achieve the best calibration (global temperature ≈ 0.95–1.00) whereas LayerNorm models are overconfident (global temperature > 1.05).
-
Per-species behavior. Because the split has one validation species (NCBI 10090, mouse) and one test species (7955, zebrafish), the per-species plots show one bar per split. BN and CLN are statistically tied on mouse within error bars; on zebrafish, CLN and BN remain close with LN slightly behind. The authors read this as support for CLN as a robust alternative to BN for cross-species generalization.
-
Cost–accuracy frontier. On the authors' hardware, SAGE variants dominate the Pareto frontier of AUC versus wall-clock time and VRAM; marker size in the cost figure is proportional to parameter count. CLN adds parameters relative to LN/BN but remains efficient, and GIN remains a weaker baseline.
-
Label co-occurrence is very sparse. 112 labels total; sparsity 0.89%; mean correlation 0.009; maximum correlation 0.061; minimum correlation 0.0003; mean outgoing correlation 0.999996; mean incoming correlation 0.999996. The authors interpret this as mostly independent functions with a few strongly interdependent functional modules.
Methodology in Plain English
The authors take the standard OGB ogbn-proteins protocol and vary one design choice at a time so that effects are attributable.
- Turning edges into node inputs. Each node's incident edges carry 8-dimensional evidence vectors. These are collapsed into a single node input using
mean,sum, ormax. This step is deliberately kept simple so its contribution can be measured in isolation. - Using edges inside message passing. Each 8-D edge vector is reduced to a single non-negative scalar α (by default the
sumof its 8 channels; a learned 1-D scalarizer is also tested), and that scalar weights the messages in GraphSAGE or GIN. - Normalization variants. The authors compare LayerNorm (per node), BatchNorm (per batch), and a Conditional LayerNorm whose affine parameters are predicted from a per-species descriptor (
cln_mode=desc). Unless stated otherwise, SAGE uses hidden size 512 and 3 layers. - No-graph control. A 3-layer MLP predicts the 112 logits directly from the edge-aggregated node features, using BatchNorm/LayerNorm/CLN + LeakyReLU + dropout in each hidden layer, to isolate how much of the performance comes from message passing rather than from edge evidence alone.
- Post-hoc calibration. Temperature scaling is applied to validation logits — either one global temperature or per-label temperatures with L2 regularization toward the global value — optimized by validation negative log-likelihood. Per-label thresholds are then chosen to maximize the F-beta score on validation probabilities, falling back to ROC-based thresholds in degenerate cases. The learned calibration and thresholds are applied to test without refitting.
- Label smoothing. A label co-occurrence matrix P is built from training labels only (no leakage), rows normalized to sum to one, and logits are nudged as z' = z + λzPᵀ before thresholding.
- Reporting. Three seeds (1, 2, 3) with early stopping on validation AUC; primary metric mean ROC-AUC across the 112 labels; secondary metric micro-F1 at a fixed 0.5 threshold; plus ECE and Brier score after calibration. Hardware is an NVIDIA A800-SXM4-40GB. All runs export standardized artifacts and the scripts are released.
Why This Matters
Impact on research. Benchmarks like ogbn-proteins are dominated by leaderboard AUC, which can hide the fact that a model's actual binary decisions at a fixed threshold are poor. This paper shows that choice of edge aggregation, normalization, and post-hoc calibration can move micro-F1 from around 0.14 to around 0.80 while AUC barely changes. That reframing — and the release of standardized artifacts and per-species, per-seed results — makes it easier for other groups to compare runs fairly rather than comparing a single headline number.
Real-world applications.
- Protein function annotation: assigning Gene Ontology-style functional labels to newly sequenced proteins, especially for organisms with little training data.
- Cross-species transfer: the species-wise split (train on some species, validate on mouse 10090, test on zebrafish 7955) mirrors the realistic case of a model trained on well-annotated organisms and deployed on a poorly annotated one.
- Laboratory screening: the paper explicitly motivates cost-sensitive thresholds for operating points such as fixed precision 90% or fixed recall 80%, which matches how wet-lab follow-up experiments are budgeted.
- Calibrated decision support: reporting ECE and Brier score matters anywhere predicted probabilities feed into downstream decisions, not just rankings.
Industry relevance. The cost dimension is reported directly: parameter counts (0.650 M for SAGE+LN and SAGE+BN, 1.710 M for SAGE+CLN, 0.852 M for GIN+BN), VRAM in megabytes for MLP variants (603–1075 MB), wall-clock time, and an explicit AUC-versus-cost Pareto view. That makes the findings usable by practitioners choosing a baseline under a memory or latency budget, and the recommendations (use sum edge aggregation, calibrate per label) are cheap to adopt on top of existing pipelines.
Future Directions
-
Unfrozen edge-network MPNN. The authors state they did not include unfrozen edge-network MPNN results due to time; a frozen-gate MPNN collapses as expected. They provide a starting configuration (
--backend sparse --hid 256 --gate_hid 64 --dropout 0.1 --epochs 120 --patience 12 --lr 2e-3 --amp 0 --norm ln/cln --alpha_chunk 2000000 --freeze_edge 0) to train the edge gate end-to-end and compare to SAGE at matched parameters and VRAM across 3 seeds, reporting AUC/F1/ECE. -
Graph Transformers with edge channels. Feeding the 8-D edge evidence as attention biases or per-head gates, compared against SAGE at matched parameters and VRAM, to test whether attention helps cross-species transfer.
-
Stronger calibration. Moving beyond temperature scaling to vector scaling, class-wise Platt scaling, or per-label isotonic regression, and adding species-conditional temperatures fitted on the validation species and applied to the test species, with micro/macro F1, ECE and reliability plots reported.
-
Cost-sensitive thresholds and learned label graphs. Optimizing thresholds for target operating points (for example fixed precision 90% or fixed recall 80%) with per-label-family precision–recall trade-offs, and replacing the heuristic matrix P with a learned label–label graph trained on validation only, augmented with GO hierarchy priors and per-label λ, comparing logit-space versus probability-space smoothing.
Target Audience
Practitioners and researchers who need a solid, reproducible baseline on ogbn-proteins or on similar edge-rich, multi-label, cross-domain graph benchmarks; PyTorch Geometric users looking for concrete guidance on edge-feature aggregation and normalization choices; and anyone studying calibration and threshold selection for multi-label GNN outputs. Readers interested in a novel architecture will not find one here — the value is in the systematic measurement, the reported compute costs, and the released artifacts.
Authors’ abstract
We present reproducible, edge-aware baselines for ogbn-proteins in PyTorch Geometric (PyG). We study two system choices that dominate practice: (i) how 8-dimensional edge evidence is aggregated into node inputs, and (ii) how edges are used inside message passing. Our strongest baseline is GraphSAGE with sum-based edge-to-node features. We compare LayerNorm (LN), BatchNorm (BN), and a species-aware Conditional LayerNorm (CLN), and report compute cost (time, VRAM, parameters) together with accuracy (ROC-AUC) and decision quality. In our primary experimental setup (hidden size 512, 3 layers, 3 seeds), sum consistently beats mean and max; BN attains the best AUC, while CLN matches the AUC frontier with better thresholded F1. Finally, post-hoc per-label temperature scaling plus per-label thresholds substantially improves micro-F1 and expected calibration error (ECE) with negligible AUC change, and light label-correlation smoothing yields small additional gains. We release standardized artifacts and scripts used for all of the runs presented in the paper.