Research
AI for pRedicting Exacerbations in KIDs with aSthma (AIRE-KIDS)
Overview Research area: Clinical machine learning / pediatric respiratory medicine — predicting asthma exacerbations from electronic medical records. Technical level: Intermediate. The clinical framin

- arXiv
- 2511.01018
- Published
- 2025-11-02
- Authors
- Hui-Lee Ooi, Nicholas Mitsakakis, Margerie Huet Dastarac, Roger Zemek, Amy C. Plint, Jeff Gilchrist, Khaled El Emam, Dhenuka Radhakrishnan
AI summary
Overview
Research area: Clinical machine learning / pediatric respiratory medicine — predicting asthma exacerbations from electronic medical records.
Technical level: Intermediate. The clinical framing is accessible, but the methods involve boosted-tree models, large language models, calibration, and SHAP-based feature attribution.
Scope: The paper develops and validates machine learning models that predict repeat severe asthma exacerbations in children who have already visited an emergency department for asthma, using hospital EMR data linked to environmental and neighbourhood information.
What This Paper Is About
Many children with asthma return to the emergency department or are hospitalized again, even though these repeat events are often preventable if the right children are identified and referred for comprehensive preventative care. The goal of AIRE-KIDS is to build machine learning models that flag which children, after a first asthma-related ED visit, are most likely to have another severe exacerbation — so that limited preventative resources can be directed toward them.
Key Contributions
- Development of machine learning models for predicting repeat severe asthma exacerbations (future asthma-related ED visits or hospital admissions) in children with a prior asthma ED visit at a tertiary care children's hospital.
- Construction of two model variants — AIRE-KIDS_ED and AIRE-KIDS_HOSP — targeting different downstream outcomes, each with an identified set of most predictive features.
- A training dataset that links Epic EMR data from the Children's Hospital of Eastern Ontario (CHEO) with environmental pollutant exposure and neighbourhood marginalization information, going beyond purely clinical variables.
- A comparison of boosted-tree methods against three open-source large language model approaches, with models tuned, calibrated, and then externally validated on a separate later-period dataset.
Main Findings
- Best-performing model: The LGBM (boosted trees) model performed best overall; the abstract reports that it outperformed the other model families tried.
- AIRE-KIDS_ED predictive features: The most predictive features in the final ED model were prior asthma ED visit, the Canadian triage acuity scale, medical complexity, food allergy, prior ED visits for non-asthma respiratory diagnoses, and age.
- Reported ED model performance: AUC of 0.712 and F1 score of 0.51 on the validation dataset.
- Comparison to existing practice: The abstract describes this as a nontrivial improvement over the current decision rule, which has an F1 of 0.334.
- AIRE-KIDS_HOSP predictive features: The hospital-admission model's most predictive features were medical complexity, prior asthma ED visit, average wait time in the ED, the pediatric respiratory assessment measure score at triage, and food allergy. The abstract does not report performance metrics for this model.
- Temporal validation design: Models trained on pre-COVID-19 data were validated on a separate post-COVID-19 dataset, distinguishing this from a simple random split.
Methodology in Plain English
The researchers took records from a children's hospital for children who had come to the emergency department for asthma. They used data from a pre-pandemic period (February 2017 to February 2019, 2,716 children) to teach the models, then checked how well those models worked on a much later period (July 2022 to April 2023, 1,237 children) — a deliberate test of whether patterns learned before COVID-19 still hold afterwards.
Beyond routine medical record variables, they enriched each record with information about environmental pollutant exposure and neighbourhood marginalization, on the theory that these factors shape asthma risk.
They tried two broad families of models: boosted decision trees (LGBM and XGB), which are standard for tabular clinical data, and three open-source large language models (DistilGPT2, Llama 3.2 1B, and Llama-8b-UltraMedical). All models were tuned and calibrated. Performance was judged with AUC and F1 scores, and SHAP values were used to reveal which individual features drove each model's predictions — the source of the feature lists reported for both model variants.
Why This Matters
- Research impact: It adds evidence on how far machine learning can push asthma risk prediction beyond existing clinical decision rules, and it tests whether models survive a temporal shift as disruptive as the COVID-19 period.
- Real-world applications:
- Flagging children at a post-ED visit who should be referred to comprehensive asthma care.
- Identifying modifiable or contextual risk factors (medical complexity, food allergy, prior respiratory visits) that clinicians could act on during follow-up.
- Supporting triage and resource allocation in pediatric emergency departments.
- Informing how environmental and neighbourhood-level data can be folded into hospital prediction tools.
- Industry relevance: The comparison of boosted trees against open-source LLMs is directly relevant to health systems and vendors deciding whether small, locally deployable language models are worth the complexity for structured clinical prediction, or whether conventional tabular methods remain the pragmatic choice.
Future Directions
- Determining whether predictive performance can be improved — the reported AUC of 0.712 and F1 of 0.51 leave substantial room, and the abstract does not describe how the models compare to alternative thresholds or operational cutoffs.
- Reporting and interrogating the performance of the AIRE-KIDS_HOSP model, whose feature set is given but whose metrics are not stated in the abstract.
- Examining how well the models generalize beyond a single tertiary care centre, since both training and validation data come from CHEO.
- Assessing the clinical utility of the models in practice — whether acting on their predictions actually reduces repeat exacerbations, which the abstract's accuracy metrics cannot answer.
- Understanding why the LLM approaches did not come out ahead of boosted trees, and whether different prompting, fine-tuning, or model sizes would change that.
Target Audience
Pediatric pulmonologists and emergency medicine clinicians, clinical informatics and health-data-science teams, machine learning researchers working on tabular clinical prediction, and health system planners designing referral pathways for chronic childhood conditions. The paper is most useful to readers with some familiarity with predictive modelling metrics (AUC, F1) and feature importance methods, though the clinical problem is described clearly enough for a general medical audience.
Authors’ abstract
Recurrent exacerbations remain a common yet preventable outcome for many children with asthma. Machine learning (ML) algorithms using electronic medical records (EMR) could allow accurate identification of children at risk for exacerbations and facilitate referral for preventative comprehensive care to avoid this morbidity. We developed ML algorithms to predict repeat severe exacerbations (i.e. asthma-related emergency department (ED) visits or future hospital admissions) for children with a prior asthma ED visit at a tertiary care children's hospital. Retrospective pre-COVID19 (Feb 2017 - Feb 2019, N=2716) Epic EMR data from the Children's Hospital of Eastern Ontario (CHEO) linked with environmental pollutant exposure and neighbourhood marginalization information was used to train various ML models. We used boosted trees (LGBM, XGB) and 3 open-source large language model (LLM) approaches (DistilGPT2, Llama 3.2 1B and Llama-8b-UltraMedical). Models were tuned and calibrated then validated in a second retrospective post-COVID19 dataset (Jul 2022 - Apr 2023, N=1237) from CHEO. Models were compared using the area under the curve (AUC) and F1 scores, with SHAP values used to determine the most predictive features. The LGBM ML model performed best with the most predictive features in the final AIRE-KIDS_ED model including prior asthma ED visit, the Canadian triage acuity scale, medical complexity, food allergy, prior ED visits for non-asthma respiratory diagnoses, and age for an AUC of 0.712, and F1 score of 0.51. This is a nontrivial improvement over the current decision rule which has F1=0.334. While the most predictive features in the AIRE-KIDS_HOSP model included medical complexity, prior asthma ED visit, average wait time in the ED, the pediatric respiratory assessment measure score at triage and food allergy.