Research
Learning from Historical Activations in Graph Neural Networks
Overview Research area: Graph machine learning, specifically graph neural network (GNN) readout and pooling layers for graph-level prediction. Technical level: Intermediate. The paper assumes familiar

- arXiv
- 2601.01123
- Published
- 2026-01-03
- Authors
- Yaniv Galron, Hadar Sinai, Haggai Maron, Moshe Eliasof
AI summary
Overview
Research area: Graph machine learning, specifically graph neural network (GNN) readout and pooling layers for graph-level prediction.
Technical level: Intermediate. The paper assumes familiarity with GNN message passing, attention mechanisms, and standard graph benchmarks (TU, OGB), but the core idea is explained with clear notation and worked equations.
Scope in one sentence: The paper proposes HistoGraph, a two-stage attention-based pooling layer that aggregates a node's representations from all GNN layers ("historical activations") rather than only the final layer, and shows it improves graph classification, node classification, and link prediction.
What This Paper Is About
Most GNN pooling and classification schemes discard the intermediate node representations produced during the forward pass and read out only the last layer's features. This wastes information, especially in deep GNNs where a node's representation can change substantially across layers and where over-smoothing makes late-layer embeddings of different nodes nearly indistinguishable. The authors introduce HistoGraph, a drop-in aggregation layer that treats each node's per-layer embeddings as a sequence, learns which layers matter, then applies node-wise self-attention at readout.
Key Contributions
- A "self-reflective" architectural paradigm for GNNs that uses the full trajectory of node embeddings across layers as a signal for the final prediction, rather than only the final-layer features.
- HistoGraph, a two-stage self-attention mechanism that disentangles (i) the layer-wise evolution of each node's embedding and (ii) the spatial aggregation of node features across the graph.
- An empirical demonstration on graph-level classification (TU and OGB), node classification, and link prediction (OGBL-COLLAB) showing consistent improvements over strong GNN and pooling baselines.
- A demonstration that HistoGraph can be used as a lightweight post-processing head on a frozen, pretrained GNN, improving models already trained with standard pooling.
Main Findings
- TU graph classification: With a 5-layer GIN backbone, HistoGraph achieves state-of-the-art performance on 5 of 7 datasets: imdb-b 87.2% ±1.7, imdb-m 61.9% ±5.5, mutag 97.9% ±3.5, proteins 97.8% ±0.4, and nci1 85.9% ±1.8. It scores 79.1% ±4.8 on ptc (marginally behind DKEPool's 79.6%) and 93.4% ±0.9 on rdt-b.
- Margins over the second-best method: The largest gains are on proteins (+16.6%), imdb-b (+6.3%), and imdb-m (+5.6%).
- OGB molecular property prediction: With a 3-layer GCN backbone, HistoGraph achieves the top ROC-AUC on 3 of 4 datasets: molbbbp 72.02% ±1.46, moltox21 77.49% ±0.70, and toxcast 66.35% ±0.80. Margins over second-best are +2.29% on molbbbp (vs. DKEPool), +0.91% on toxcast (vs. GMT), and +0.19% on moltox21 (vs. GMT). On molhiv, DKEPool leads with 78.65% while HistoGraph reaches 77.81% ±0.89.
- Robustness to depth (over-smoothing): On node classification, plain GCN degrades sharply with depth while GCN + HistoGraph stays stable. On Cora, GCN drops from 81.1% at 2 layers to 28.7% at 64 layers, whereas GCN + HistoGraph moves from 81.3% to 77.5%. On Citeseer, GCN falls from 70.8% to 20.0% while GCN + HistoGraph falls only from 70.9% to 63.4%. On Pubmed, GCN drops from 79.0% to 35.3% while GCN + HistoGraph rises from 78.9% to 79.3%.
- Post-processing on frozen backbones: Applied as a frozen auxiliary head (FT) on pretrained GINs with mean pooling, HistoGraph raises imdb-m accuracy from 54.7% (MeanPool) to 67.3% (FT), and imdb-b from 76.0% to 94.0% (both FT and full fine-tuning). On proteins it reaches 97.3% (FT and Full FT) versus 75.9% for MeanPool; on ptc, full fine-tuning gives the best score at 97.1% versus 77.1% for MeanPool.
- Ablation (proteins, 97.80% ±0.40 for full HistoGraph): Removing division-by-sum normalization drops accuracy to 74.45% ±6.28, removing layer-wise attention gives 78.61% ±4.82, and removing node-wise attention gives 80.78% ±7.71. DKEPool is listed at 81.20% ±3.80 on this dataset.
- Signed normalization as adaptive filtering: Unlike softmax, the division-by-sum normalization allows negative layer weights, letting the model implement low-pass (uniform average), high-pass (first difference), or general FIR-style filters over the layer trajectory. The authors illustrate this on a barbell graph, where a trained GCN fails to reproduce a node-feature gradient target that HistoGraph captures.
- Runtime: End-to-end HistoGraph is costlier than MeanPool during training (e.g., 60.34s vs. 41.27s per epoch for 32 layers on molhiv) but remains scalable; the frozen-head variant reduces overhead, and at inference HistoGraph adds negligible overhead over MeanPool. It is reported as significantly faster than GMT in almost all cases.
- Complexity: Per-graph cost is O(NLD + N²D) = O(N(L+N)D) with O(L + N²) memory, versus O(L²N²D) for a naive joint node–layer attention and O(LN²D) for a graph transformer stacking L attention layers.
Methodology in Plain English
The authors start from a standard pretrained-style GNN backbone that produces a hidden embedding for every node at every layer. Instead of throwing away all but the last layer, they stack these embeddings into a per-node "history" tensor and add fixed sinusoidal positional encodings so the model knows which layer each entry came from.
Two attention stages then run in sequence. The first, layer-wise attention, uses only the final-layer embedding as a query and attends over all previous layers, producing a single blended embedding per node. This attention is averaged across nodes and normalized by dividing by the sum of scores rather than applying softmax, which permits negative weights and lets the layer mixture act as an adaptive filter over the node's trajectory. The second, node-wise multi-head self-attention operates once at the readout stage, letting nodes exchange information across the graph before a final mean over nodes produces the graph descriptor. Because the node-wise attention runs only once (not at every message-passing layer), it avoids adding to over-smoothing during propagation.
The layer can be trained end-to-end with the backbone or used as a head on a frozen backbone, in which case the N × L × D activations are cached in a single forward pass and no gradients flow through the GNN layers. Experiments compare against a broad set of pooling methods (Set2Set, SortPool, SAGPool, TopKPool, ASAP, DiffPool, MinCutPool, HaarPool, StructPool, EdgePool, GMT, SOPool, HAP, PAS, GMN, DKEPool, JKNet, and kernel baselines), using 5-layer GIN backbones on seven TU datasets and 3-layer GCN backbones on four OGB datasets, plus node-classification datasets and link prediction. Hardware used was NVIDIA L40, A100, and GeForce RTX 4090 GPUs with PyTorch, PyTorch Geometric, and Weights and Biases.
Why This Matters
Impact on research: The paper reframes intermediate GNN activations as a first-class signal for readout rather than a byproduct. It offers a general, architecture-agnostic layer that can be bolted onto existing GNNs, gives a formal proposition showing how layer-wise weighting preserves distinguishability under over-smoothing, and connects learned layer weights to filter theory. This provides a template for future work on multi-scale readout and depth-robust GNN design.
Real-world applications (domains the paper references):
- Molecular property prediction and drug discovery, where graph classification on molecules is the target task.
- Social network analysis, where graphs are large and structure is complex.
- Recommendation systems built on user–item interaction graphs.
- General graph-level classification pipelines (e.g., protein and bioinformatics datasets such as proteins, ptc, mutag, nci1) that need a fixed-size graph descriptor.
Industry relevance: Because HistoGraph works as a lightweight post-processing head on a frozen pretrained GNN, teams can improve already-deployed graph models without retraining the encoder — attractive where retraining is expensive or data is limited (the paper notes low-resource and few-shot regimes). Its modest runtime overhead and a single O(N²D) attention pass also keep it practical relative to stacking full graph-transformer layers.
Future Directions
- Scaling HistoGraph beyond small-to-medium graphs: the authors state the current design is best suited to this regime and point to sparse attention or hierarchical coarsening as promising extensions (discussed further in Appendix G).
- Reducing the quadratic O(N²D) node-wise attention cost so the layer applies to very large graphs.
- More extensive validation across task types — graph classification, node classification, and link prediction (OGBL-COLLAB) — and across additional backbones beyond GIN and GCN.
- Further understanding and regularizing the learned signed layer weights, since removing the division-by-sum normalization causes the largest ablation drop and unbounded weights destabilize training.
Target Audience
Researchers and practitioners in graph machine learning who work on GNN architectures, graph pooling, or readout layers, and who want a drop-in component that improves deep GNNs and frozen pretrained models. It is also useful for applied scientists in cheminformatics, social network analysis, and recommendation who need stronger graph-level representations without redesigning their backbone. Readers should already be comfortable with attention and message passing; the math is intermediate rather than introductory.
Authors’ abstract
Graph Neural Networks (GNNs) have demonstrated remarkable success in various domains such as social networks, molecular chemistry, and more. A crucial component of GNNs is the pooling procedure, in which the node features calculated by the model are combined to form an informative final descriptor to be used for the downstream task. However, previous graph pooling schemes rely on the last GNN layer features as an input to the pooling or classifier layers, potentially under-utilizing important activations of previous layers produced during the forward pass of the model, which we regard as historical graph activations. This gap is particularly pronounced in cases where a node's representation can shift significantly over the course of many graph neural layers, and worsened by graph-specific challenges such as over-smoothing in deep architectures. To bridge this gap, we introduce HISTOGRAPH, a novel two-stage attention-based final aggregation layer that first applies a unified layer-wise attention over intermediate activations, followed by node-wise attention. By modeling the evolution of node representations across layers, our HISTOGRAPH leverages both the activation history of nodes and the graph structure to refine features used for final prediction. Empirical results on multiple graph classification benchmarks demonstrate that HISTOGRAPH offers strong performance that consistently improves traditional techniques, with particularly strong robustness in deep GNNs.