Research
PhiPlot: A Web-Based Interactive EDA Environment for Atmospherically Relevant Molecules
PhiPlot: A Web-Based Interactive EDA Environment for Atmospherically Relevant Molecules Overview Research area: Human-Computer Interaction / Scientific Visualization, applied to computational atmosphe

- arXiv
- 2603.11751
- Published
- 2026-03-12
- Authors
- Matias Loukojärvi, Ananth Mahadevan, Katsiaryna Haitsiukevich, Kai Puolamäki
AI summary
PhiPlot: A Web-Based Interactive EDA Environment for Atmospherically Relevant MoleculesOverview
Research area: Human-Computer Interaction / Scientific Visualization, applied to computational atmospheric chemistry (specifically atmospheric aerosol formation).
Technical level: Intermediate. The paper is a systems-and-evaluation demo paper; it assumes some familiarity with dimensionality reduction and molecular fingerprints but explains the pipeline accessibly.
One-sentence scope: The paper introduces PhiPlot, a web-based exploratory data analysis environment that connects to a curated store of molecular databases so researchers can summarise, filter, cluster, and interactively embed atmospherically relevant molecules using knowledge-guided dimensionality reduction.
What This Paper Is About
Computational chemistry has generated large, high-dimensional datasets of molecular structures and properties, but it remains an open problem to identify which combinations of properties trigger processes such as particle formation in the atmosphere. The authors argue that datasets in this domain are also incomplete, because quantum-level accurate simulation for every potential molecule is computationally prohibitive. PhiPlot is a web application that lowers the barrier to exploring these datasets by combining summary statistics, clustering, static embeddings, and real-time interactive constraint-based embeddings in a single interface connected to an evolving collection of molecular databases.
Key Contributions
- Application of existing data analysis methods to a new domain. The authors apply knowledge-based dimensionality reduction to atmospherically relevant molecules, using methods such as Constrained Kernel PCA (cKPCA) and Least Square Projections (LSP).
- A web-based environment integrated with a curated collection of molecular databases. PhiPlot connects directly to a MongoDB-backed document-oriented NoSQL data store and requires no user-side installation or configuration.
- A small user study verifying the application's utility. Six subjects with expertise in atmospheric chemistry or computer science evaluated the tool via Likert-scale statements and open questions.
- A comparative positioning against prior tools. Table 1 contrasts PhiPlot against InVis/InVis 2.0, Solvent Surfer, and χiplot against the four design goals; the authors state that, to the best of their knowledge, no prior method explicitly addresses all of the design goals.
Main Findings
- Interactive constraints are needed to make static projections meaningful. The user study highlighted that interactions are required to make the initial static 2D projection more meaningful (statement 3 of the survey).
- Domain experts gained insights even when already familiar with the data. The survey answers show that domain experts leveraged interactions to embed domain knowledge (statements 4 and 6).
- Non-chemistry experts also found it useful. Computer science experts with only basic familiarity with chemistry found the application insightful (statement 2) and helpful for fast data exploration (statements 1 and 5).
- Speed was rated positively by all experts. All experts agreed that the implementation was fast (statement 5). The paper reports that a cKPCA embedding update takes less than a second for N ≲ 500 embedded points, enabling near real-time dragging of control points.
- Qualitative praise from chemistry experts. Atmospheric chemistry domain experts described PhiPlot as "easy and intuitive to use" and a "fast way of looking at the general statistics and structure of the dataset."
- Qualitative praise from computer science experts. Computer science experts noted that the "UI is fast and responsive" and that there is a "good selection of embeddings out-of-the-box."
- Requested improvements (chemistry experts). More options for the datapoint search function, and the ability to "select the group of molecules looked by [its] precursors."
- Requested improvements (computer science experts). The ability to "filter and observe the number of certain functional groups."
- Comparison with prior work. InVis and InVis 2.0 are non-web-based and require users to load their own data, applying only dimensionality reduction without built-in summary statistics or clustering. Solvent Surfer is web-based with an integrated solvent dataset but lacks comprehensive EDA and molecular representation functionality and does not support real-time interactive embedding updates. χiplot offers static embeddings, EDA plots, and clustering, but does not support molecular fingerprinting or interactive embedding methods.
Methodology in Plain English
Data representation. Molecules are stored as SMILES strings encoding their 2D structures, alongside properties from experiments or simulations. The application converts these strings into numerical feature vectors called molecular fingerprints, which check for predefined substructures or enumerate and hash molecular fragments. The paper names RDKit as a general-purpose fingerprinting library and ATMOMACCS as a technique developed specifically for atmospheric molecules.
Dimensionality reduction. Fingerprints are high-dimensional, sparse bit or integer vectors. To visualise them, they are projected into a 2D space, producing what the paper calls an embedding. Typical methods are unsupervised and produce only static embeddings; PhiPlot adds semi-supervised constrained methods.
Interactive embedding. Constrained Kernel PCA (cKPCA) is described as a fast semi-supervised method allowing real-time interactivity. Users add domain knowledge through control points (points dragged to a desired position) and link constraints. Must-links assert that two points should be closer; cannot-links push two points apart. Link-constraint strength is adjustable with a slider.
Interface and architecture. The application has three main views displaying progressively increasing levels of detail, following Shneiderman's "overview, zoom, filter, details-on-demand" principle. The data store is MongoDB-backed, and summary computations use database-native aggregation pipelines to avoid extra data transfer. The interface is written in Python with Holoviz Panel and HoloViz-maintained libraries, with Bokeh as the plotting back-end. It is containerised with Docker, deployed with OpenShift, and the prototype (database backend plus visualisation tools) is hosted on the CSC — IT Centre for Science server.
Evaluation metrics built into the interface. Clustering quality is shown via Silhouette, Calinski-Harabasz, and Davis-Bouldin scores. Embedding quality is shown via trustworthiness, KNN preservation, Shepard correlation, and stress.
Design goals guiding the build. DG1 data accessibility (several molecular databases in one place), DG2 application accessibility (no installation or configuration), DG3 flexible data exploration (toggle between clustering/dimensionality reduction and high-level summary statistics), and DG4 interactive data analysis (constraint embeddings for human-in-the-loop alignment with the expert's mental model).
User study protocol. Sessions lasted approximately 30 minutes, in person or via Zoom. Subjects received a brief introduction and a short tutorial, then completed tasks related to each view and main feature, with the interviewer guiding them if they felt lost. They then filled out a survey ranking six statements on a 5-point Likert scale (1 = disagree, 5 = agree) and answered two open questions about benefits compared to previous tools and workflows, and the most important improvements. The study had a total of N = 6 subjects: 4 with atmospheric chemistry expertise (2 senior researchers, 2 doctoral students) with varying levels of computer science expertise, and 2 computer science doctoral students with at least some familiarity with atmospheric chemistry. Categorisation in the results is based on main domain expertise.
Why This Matters
Impact on research. The paper positions PhiPlot as bridging the gap between raw high-dimensional molecular data and scientific insight, supporting hypothesis generation and informed subsetting of molecules. The authors note that the user study confirms the design goals were overall met.
Real-world applications:
- Exploring molecular datasets for atmospheric aerosol formation, where the specific property combinations that trigger particle formation remain an open problem.
- Chemical data exploration in the broader sense of navigating large computational chemistry datasets, given that quantum-accurate simulation for every candidate molecule is computationally prohibitive.
- Domain-expert-guided analysis workflows, where a chemist's prior knowledge is injected directly into a visualisation through control points and link constraints rather than through code.
- Rapid assessment of existing datasets, since the summary statistics view lets users inspect available covariates with filters applied without any local setup.
Industry relevance. The paper does not report deployment in an industrial or commercial setting; it describes a publicly available prototype hosted on the CSC — IT Centre for Science server, with source code available at https://github.com/edahelsinki/phiplot. Any industrial relevance would follow from the general applicability of integrated EDA and interactive dimensionality reduction, not from anything demonstrated in the paper.
Future Directions
- Advanced filtering features based on chemical structure, which the authors state can simplify the interactions.
- Extension to explore and develop machine learning models trained to predict molecular properties.
- More options for the datapoint search function, as requested by the chemistry experts in the user study.
- Selecting groups of molecules by their precursors, another chemistry expert request.
- Filtering and observing the number of certain functional groups, as requested by the computer science experts.
Target Audience
This paper is most useful for researchers and practitioners in human-computer interaction and visual analytics who are interested in domain-knowledge-guided dimensionality reduction, and for computational and atmospheric chemists who need to explore high-dimensional molecular datasets. It is also relevant to developers building browser-based scientific tooling who want an example of combining a MongoDB data store, HoloViz Panel, and Bokeh with interactive constrained embeddings. Readers looking for a rigorous quantitative benchmark or a large-scale evaluation will not find one here: the user study has N = 6 subjects, and the paper itself describes it as small.
Authors’ abstract
Advances in computational chemistry have produced high-dimensional datasets on atmospherically relevant molecules. To aid exploration of such datasets, particularly for the study of atmospheric aerosol formation, we introduce PhiPlot: a web-based environment for interactive exploration and knowledge-based dimensionality reduction. The integration of visualisation, clustering, and domain knowledge-guided embedding refinement enables the discovery of patterns in the data and supports hypothesis generation. The application connects to an existing, evolving collection of molecular databases, offering an accessible interface for data-driven research in atmospheric chemistry.