Research
AutocleanEEG ICVision: Automated ICA Artifact Classification Using Vision-Language AI
AutocleanEEG ICVision: Automated ICA Artifact Classification Using Vision-Language AI Overview Research area: Neurophysiology / EEG signal processing, applying multimodal vision-language AI (computer

- arXiv
- 2512.00194
- Published
- 2025-11-28
- Authors
- Zag ElSayed, Grace Westerkamp, Gavin Gammoh, Yanchen Liu, Peyton Siekierski, Craig Erickson, Ernest Pedapati
AI summary
AutocleanEEG ICVision: Automated ICA Artifact Classification Using Vision-Language AIOverview
Research area: Neurophysiology / EEG signal processing, applying multimodal vision-language AI (computer vision plus natural language reasoning) to independent component analysis (ICA) artifact classification.
Technical level: Intermediate. The EEG/ICA and vision-language model concepts are explained clearly enough for readers with some background, but the paper assumes familiarity with terms such as ICA, topography, power spectral density, and ERPs.
Scope: The paper describes and evaluates a system, ICVision, that classifies EEG ICA components by looking at the same four-panel diagnostic images that human experts read, and that returns a label, a confidence score, and a natural-language explanation.
What This Paper Is About
EEG recordings capture brain activity but also pick up muscle activity, eye movements, heartbeats, and electrical interference. Separating brain from artifact is often done with Independent Component Analysis (ICA), which decomposes a recording into independent components but does not label them; a trained human must visually inspect diagnostic plots for every component, which does not scale as datasets grow. ICVision's goal is to automate that expert visual judgment using a multimodal large language model that "sees" the plots and explains its decision in plain language.
Key Contributions
-
A vision-based (rather than feature-based) ICA classifier. ICVision sends full four-panel ICA dashboard images — topographic map, time series, power spectral density, and ERP image — to a multimodal model (OpenAI GPT-4 Vision API, with GPT-4.1 multimodal also described as the chosen model) instead of relying on handcrafted numerical features like ICLabel.
-
Explainable, auditable output. Every classification returns a class label, a confidence score between 0 and 1, and a 30–70 word natural-language justification written in clinician-style reasoning, which the authors frame as the first implementation of AI-agent visual cognition in neurophysiology.
-
A full open-source preprocessing pipeline with auto-exclusion. ICVision is a core module of the open-source EEG Autoclean platform, producing a defined output catalog (Results.csv, CleanedRawSet.fif, ReportAllComp.pdf, Summary.txt) and applying confidence-based filtering to auto-reject artifacts while preserving or flagging uncertain components.
-
An expert-annotated benchmark of real ICA dashboards. Performance was assessed against human expert labels on 3,168 ICA components drawn from 124 EEG datasets, with a direct comparison against MNE-ICLabel.
Main Findings
-
Agreement with experts: ICVision achieved a Cohen's kappa of 0.677 against human expert labels, versus kappa = 0.661 for MNE-ICLabel on the same dataset, with exact agreement with expert labels in 59% of components.
-
Interpretability ratings: Over 97% of outputs were rated by expert reviewers as interpretable and actionable; explanations were rated excellent in 85% of reviewed cases and acceptable in 12%.
-
Agreement stratification: 67.6% of components were unanimously classified across ICVision, ICLabel, and human labels; 13.5% showed mixed agreement; 18.9% were labeled similarly by ICVision and ICLabel but diverged from human labels.
-
Class distribution: Muscle was the most frequently assigned class (35.5%), followed by brain (30.6%), eye (16.9%), channel noise (16.1%), and heart (0.8%). "Other noise" was not directly assigned in this implementation, which the authors attribute to it being merged with adjacent artifact classes through visual context interpretation.
-
Signal preservation: In one illustrated case, ICVision retained a frontal component that ICLabel rejected as eye, preserving a clear alpha peak at 8.2 Hz that was lost in the ICLabel-cleaned data.
-
Scale of annotation effort: Manual labeling was estimated at roughly 30 seconds per component, reaching 100 hours for the dataset. A total of 3,168 components (from 10,967 images processed) were independently labeled by an EEG analyst with over 7 years of ICA interpretation experience, with discrepancies resolved by consensus review.
-
Cost and throughput: Batch size was 10 components per request with up to 4 concurrent API calls via Python asyncio, at an average cost of approximately $0.002 per component using gpt-4.1-mini (roughly 50 cents for 128 components).
-
Reported figures differ across sections: The abstract and Section III-B report kappa = 0.677; Section I-E separately states 95% agreement with expert consensus; the system is described as having six canonical categories in the abstract and Section III-A but "seven class labels" / seven predefined categories in Sections I-C and II-D. These inconsistencies appear in the paper itself.
Methodology in Plain English
Raw EEG data (clinical and cognitive recordings from 128-channel systems using standard 10-20 montages) were preprocessed with bandpass filtering at 1–80 Hz, adaptive notch filters for line noise, and channel rejection and interpolation using MNE-Python and EEGLAB. ICA was then run with EEGLAB's ICA or MNE's FastICA depending on platform, yielding 20–40 components per subject.
For each component, the team generated a 512×512 four-panel image mimicking an expert's diagnostic screen: a topographic map showing spatial distribution across the scalp in microvolts, a time series showing activation over a 2.5-second segment, a power spectral density plot in dB-scaled power across 1–80 Hz, and an ERP image showing amplitude across data epochs. These .png dashboards, named by subject and component index, were submitted to the vision-language model alongside a task-specific prompt instructing the model to act as a neurologist, assign one of the artifact or brain categories, give a confidence score, and explain briefly.
The model processes the whole image at once, interpreting spatial, temporal, and spectral cues together. Its structured output (label, confidence, explanation) feeds a rule-based filtering step: artifact-labeled components with confidence at or below 0.80 are auto-marked for rejection, and brain-labeled components are retained unless confidence exceeds 0.40, in which case they are flagged for manual review. Rejected components are zeroed out and the remaining ICA solution is projected back into the cleaned EEG. A consistency log was kept to make reprocessing reproducible. The authors note the model is stateless and deterministic under fixed prompts.
Why This Matters
Impact on research: ICA-based cleaning is a reproducibility bottleneck in EEG science — decisions are subjective, undocumented, and lost when personnel leave. ICVision reframes the task as visual-linguistic reasoning, making cleaning decisions traceable, consistent across sessions and sites, and auditable, while separating performance from handcrafted feature engineering.
Real-world applications:
- Clinical neurodiagnostics, where cleaning decisions must be explainable and traceable for regulatory compliance and stakeholder trust.
- Brain-computer interface (BCI) pipelines, where preserving genuine neural spectral content such as alpha peaks affects downstream decoding.
- Biomarker detection in cognitive neuroscience, where over-rejecting borderline neural components can distort research outcomes.
- Training and knowledge transfer, since the natural-language explanations can teach trainees and preserve institutional EEG interpretation logic.
Industry relevance: By reducing per-component cost to approximately $0.002 and supporting batched, asynchronous processing compatible with MNE-Python 1.4+, EEGLAB 2023+, and both .set and .fif formats (Python 3.10, MATLAB R2023a), the approach targets hospital-grade, multi-site workflows rather than single-lab scripts, and points toward a broader class of explainable AI agents in neuroscience.
Future Directions
- Fine-tuning and specialized backends: The authors chose GPT-4.1 multimodal over a fine-tuned model because of its strong baseline, but explicitly note fine-tuning for specialized tasks or cohorts, and integration of domain-tuned biomedical models, as future options.
- Improving the classifier's weakest agreement areas: The 18.9% of components where ICVision and ICLabel agreed but diverged from human labels remains an open question about where automated consensus and expert judgment part ways.
- Restoring the missing class: "Other noise" was never directly assigned in this implementation, raising the question of how to recover or redefine that category.
- Extension beyond ICA cleaning: Applying the same vision-language agent approach to pediatric, clinical, and real-time domains, and to other forms of neurological signal interpretation, is proposed as a natural extension.
Target Audience
EEG researchers and clinicians who preprocess recordings and need repeatable artifact handling; neurophysiologists and EEG technicians interested in how their visual reasoning can be captured and scaled; brain-computer interface and cognitive neuroscience engineers concerned with signal preservation; machine learning practitioners interested in vision-language models applied to scientific visualization rather than natural images; and clinical informatics or regulatory stakeholders evaluating explainable, auditable automation for medical data pipelines.
Authors’ abstract
We introduce EEG Autoclean Vision Language AI (ICVision) a first-of-its-kind system that emulates expert-level EEG ICA component classification through AI-agent vision and natural language reasoning. Unlike conventional classifiers such as ICLabel, which rely on handcrafted features, ICVision directly interprets ICA dashboard visualizations topography, time series, power spectra, and ERP plots, using a multimodal large language model (GPT-4 Vision). This allows the AI to see and explain EEG components the way trained neurologists do, making it the first scientific implementation of AI-agent visual cognition in neurophysiology. ICVision classifies each component into one of six canonical categories (brain, eye, heart, muscle, channel noise, and other noise), returning both a confidence score and a human-like explanation. Evaluated on 3,168 ICA components from 124 EEG datasets, ICVision achieved k = 0.677 agreement with expert consensus, surpassing MNE ICLabel, while also preserving clinically relevant brain signals in ambiguous cases. Over 97% of its outputs were rated as interpretable and actionable by expert reviewers. As a core module of the open-source EEG Autoclean platform, ICVision signals a paradigm shift in scientific AI, where models do not just classify, but see, reason, and communicate. It opens the door to globally scalable, explainable, and reproducible EEG workflows, marking the emergence of AI agents capable of expert-level visual decision-making in brain science and beyond.