Research
Boltzmann Graph Ensemble Embeddings for Aptamer Libraries
Boltzmann Graph Ensemble Embeddings for Aptamer Libraries Overview Research area: Machine learning for biochemistry — specifically graph-based molecular representation and unsupervised community detec

- arXiv
- 2510.21980
- Published
- 2025-10-24
- Authors
- Starlika Bauskar, Jade Jiao, Narayanan Kannan, Alexander Kimm, Justin M. Baker, Matthew J. Tyler, Andrea L. Bertozzi, Anne M. Andrews
AI summary
Boltzmann Graph Ensemble Embeddings for Aptamer LibrariesOverview
Research area: Machine learning for biochemistry — specifically graph-based molecular representation and unsupervised community detection applied to aptamer selection data from SELEX experiments.
Technical level: Advanced. The paper combines exponential-family random graph models, RNA/DNA secondary-structure thermodynamics, dynamic programming with partition functions, non-negative matrix factorization, spectral clustering, and t-SNE.
One-sentence scope: The paper introduces a thermodynamically parameterized exponential-family random graph (ERGM) embedding that represents each aptamer as a Boltzmann-weighted ensemble of secondary-structure graphs rather than a single minimum-free-energy fold, and applies it to detect anomalous and promising aptamer candidates in SELEX sequencing data.
What This Paper Is About
Most machine-learning methods for biomolecules represent a molecule as a single graph — typically the minimal free energy (MFE) secondary structure — which fails to capture the thermodynamic conformational ensembles that govern weakly folded, dynamic molecules such as single-stranded DNA aptamers in solution. In SELEX, aptamer selection is driven by sequence read counts that are treated as surrogates for binding affinity, but experimental biases such as PCR amplification can make high-count sequences poor binders and low-count sequences strong binders. The goal is an embedding that averages over an aptamer's structural ensemble so that structurally similar aptamers cluster together and misleading, bias-inflated candidates can be flagged.
Key Contributions
-
Two task-aligned motif families as embedding features. The authors specify faces (typed and energy-indexed secondary-structure subgraphs) and rooted neighborhoods (isomorphic closed radius-r neighborhoods), and use their expected appearance across the Boltzmann ensemble as the feature vector for a sequence.
-
A thermodynamically parameterized ERGM embedding. The Boltzmann distribution over pseudoknot-free secondary-structure graphs is shown to be an ERGM with sufficient statistics given by motif counts and parameters θ_k = −β w_k, and the embedding is the ensemble expectation of the graph feature-count vector — an "expected bag-of-faces" or "expected bag-of-neighborhoods."
-
Application to SELEX data with community detection. The embeddings are applied to SELEX libraries against norepinephrine, where they cluster structurally similar aptamers and isolate robust communities that survive in the presence of anomalous sequences.
-
Subgraph-level explainability and anomaly analysis. A Ridge regression from the embedding to selective pressure yields interpretable face and k-mer features, which are then used to build embeddings that separate over-valued anomalies from high-performing aptamers and to recommend candidates for further testing.
Main Findings
-
Ensemble embeddings isolate robust communities. After NMF with 25 topics on the expected-neighborhood matrix (X_EN) and spectral clustering into 35 clusters, the t-SNE visualization shows circled clusters that contain no over-valued (HC-LP) anomalies. These clusters also exhibit the highest average selective pressures and counts per million.
-
Ridge coefficients identify explanatory subgraph motifs. Fitting X_EBOF w = ρ with a Ridge regressor on the 3102-dimensional expected bag-of-faces matrix, the most negative coefficients were INTERNAL:9+8:AT/TA (−0.45), ATTC (−0.13), BULGE:11:CA/GT (−0.10), CATG (−0.09), and INTERNAL:12+11:AT/TA (−0.09); the most positive were INTERNAL:6+7:AT/TG (0.50), BULGE:17:AC/TG (0.42), BULGE:11:AA/TT (0.25), TTTA (0.23), and ATTT (0.23).
-
Anomalous clusters can be discarded. Community detection on the negatively correlated features (25 NMF topics, spectral clustering into 25 clusters) flagged the 10 clusters with the highest Δ_C as anomalous. Aptamer 2, a high-count low-pressure aptamer, was identified as an over-valued anomaly, and aptamer 587 — the aptamer with the lowest selective pressure in the data set — was discarded. Most tested positive binders were not in a circled cluster.
-
Recommendations recover tested binders. Applying the same NMF and t-SNE procedure to the positively correlated features with anomalous clusters removed, confirmed good binders were located in two of the recommended clusters. Recommended clusters sit close together on the right side of the plot, indicating similarity in w_pos feature counts.
-
Data scale. From unprocessed SELEX data (two libraries; N48 rounds 9 & 13, N58 rounds 12 & 16), after removing aptamers that emerge in a later round without appearing in an earlier one, 3711 unique aptamers remained. The paper states that seven experimentally tested high-binding aptamers are used as validation, and that among the 3711, six aptamers exhibit an experimentally validated high binding affinity.
-
Embedding dimensions. The expected bag-of-faces matrix is X_EBOF ∈ ℝ^(3711 × 3102) and the expected-neighborhood matrix is X_EN ∈ ℝ^(3711 × 850), each concatenated with a k-mer embedding using k = 4.
Methodology in Plain English
The authors treat each DNA aptamer as a set of possible secondary structures rather than one structure. Starting from the nucleotide sequence, they use dynamic programming with partition functions (implemented via ViennaRNA, using DNA_Mathews_2004 nearest-neighbor parameters, a temperature of 37 °C / 310.15 K, and k_B = 1.98 × 10⁻³ kcal mol⁻¹ K⁻¹) to obtain the Boltzmann probability of each pseudoknot-free structure. Because these structures are outerplanar, they can be decomposed into faces — stack, hairpin, internal loop, bulge, and multibranch — each with an associated free energy. The authors count faces by their (type, energy) pair and also count isomorphic rooted neighborhoods of radius r = 4, then average these counts over the Boltzmann distribution to get one fixed-size vector per sequence. These expected counts are non-negative, which suits non-negative matrix factorization. They then run NMF (25 topics), spectral clustering (35 clusters chosen by sweeping 5 to 50 clusters and selecting the highest silhouette score), and t-SNE for visualization. Separately, they regress the face embedding against a "selective pressure" trend metric (the fractional change in read count between rounds, summed across libraries for aptamers in both), and use the signs of the fitted coefficients to define feature subsets for anomaly-focused clustering. To label anomalies without ground truth, they threshold counts at the 90th percentile and selective pressure at the 10th percentile, defining high-count low-pressure (HC-LP) and low-count high-pressure (LC-HP) groups.
Why This Matters
Impact on research. The work argues that single-MFE graph representations miss biologically relevant conformational flexibility, and that ensemble-weighted graph fingerprints provide a principled alternative. It also shows how partial, threshold-based labeling of biased SELEX data can be combined with interpretable embeddings to produce testable candidate lists, rather than simply picking the most abundant sequences.
Real-world applications:
- Biosensing, where aptamers are used as recognition elements (the paper cites aptamer impact on biosensing).
- Therapeutics, where aptamers serve as targeting or therapeutic molecules.
- Molecular engineering, where aptamers are designed or evolved for specific functions.
- Directing experimental follow-up in SELEX campaigns — particularly toward low-abundance candidates that would otherwise be ignored — while flagging likely PCR-amplification artifacts.
Industry relevance. Commercial and translational aptamer discovery pipelines depend on expensive post-SELEX characterization, and only a small number of candidates can be tested due to laboratory costs and manual labor. A method that prioritizes which candidates to synthesize and assay, and that explicitly down-weights amplification-biased sequences, targets that bottleneck. The authors suggest that building future initial libraries enriched for features in w_pos and depleted of features in w_neg may enhance aptamer discovery.
Future Directions
-
Learning thermodynamic parameters rather than relying on fixed ones, since fixed thermodynamic parameters are listed as a key limitation.
-
Incorporating pseudoknots, which are currently excluded from the pseudoknot-free structure set.
-
Multi-temperature ensembles, extending the single-temperature (37 °C) Boltzmann treatment.
-
Experimental validation of the recommended aptamer candidates, including the low-abundance, high-pressure sequences the method is designed to surface.
Target Audience
Researchers working at the intersection of machine learning and biochemistry — particularly those building graph-based molecular representations, working on RNA/DNA structure ensembles, or analyzing SELEX and other in-vitro evolution sequencing data. It will also be useful to computational biologists and applied mathematicians interested in exponential-family random graph models, outerplanar graph algorithms, and unsupervised anomaly detection with interpretable features. Experimental aptamer researchers may benefit from the candidate-prioritization framing, though the method itself requires familiarity with sequence data, folding thermodynamics, and matrix-factorization workflows.
Authors’ abstract
Machine-learning methods in biochemistry commonly represent molecules as graphs of pairwise intermolecular interactions for property and structure predictions. Most methods operate on a single graph, typically the minimal free energy (MFE) structure, for low-energy ensembles (conformations) representative of structures at thermodynamic equilibrium. We introduce a thermodynamically parameterized exponential-family random graph (ERGM) embedding that models molecules as Boltzmann-weighted ensembles of interaction graphs. We evaluate this embedding on SELEX datasets, where experimental biases (e.g., PCR amplification or sequencing noise) can obscure true aptamer-ligand affinity, producing anomalous candidates whose observed abundance diverges from their actual binding strength. We show that the proposed embedding enables robust community detection and subgraph-level explanations for aptamer ligand affinity, even in the presence of biased observations. This approach may be used to identify low-abundance aptamer candidates for further experimental evaluation.