Research
Incorporating data drift to perform survival analysis on credit risk
Overview Research area: Credit risk modelling — specifically survival analysis (time-to-default modelling) combined with data drift adaptation, applied to mortgage loan portfolios. Technical level: Ad

- arXiv
- 2601.20533
- Published
- 2026-01-28
- Authors
- Jianwei Peng, Stefan Lessmann
AI summary
Overview
- Research area: Credit risk modelling — specifically survival analysis (time-to-default modelling) combined with data drift adaptation, applied to mortgage loan portfolios.
- Technical level: Advanced. The paper assumes familiarity with survival analysis, discrete-time hazard models, joint modelling, landmarking, isotonic calibration, and concept-drift terminology (sudden, incremental, recurring, label drift).
- Scope: The paper proposes and evaluates a dynamic joint survival model (LMISO) for mortgage default prediction on Freddie Mac loan-level data under three simulated drift regimes plus label drift.
Note: the full paper content supplied here is truncated. It contains the abstract, practitioner summary, introduction, related work, methodology, and the beginning of the data and experimental settings section. The results section (Section 5), including all tables of numerical scores, is not included, so the specific metric values, model rankings and ablation results cannot be reported here beyond what the abstract states.
What This Paper Is About
Most credit risk survival models implicitly assume that the data-generating process is stationary — that the relationship between borrower characteristics and default risk stays the same over time. In real mortgage portfolios this assumption breaks down: borrower behaviour changes, interest rates move, housing markets cycle and policy regimes shift, all of which cause data drift and degrade model performance. The paper's goal is to design and test a survival-based default model that stays accurate and well-calibrated when the data distribution is drifting, by combining a behavioural balance-deviation marker, landmark-based discrete-time hazard modelling, time-specific baseline adjustments, and isotonic probability calibration.
Key Contributions
- First systematic drift study for survival-based credit risk models (as claimed by the authors): The paper states it is the first study to systematically investigate survival-based credit risk models under multiple concept drift scenarios — sudden, incremental and recurring drift together with label drift — within a unified experimental framework.
- A landmark-based dynamic joint modelling approach: The model integrates a longitudinal behavioural marker derived from balance dynamics (observed unpaid balance versus the scheduled amortisation path) with a discrete-time hazard formulation, so default risk can be updated dynamically as new monthly information arrives.
- A joint landmark one-hot and isotonic calibration (LMISO) strategy: Landmark one-hot encoding provides landmark-specific baseline (intercept) adjustments to capture temporal heterogeneity, while isotonic regression recalibrates predicted probabilities under drift. The authors report this combination outperforms classical survival models, machine learning methods and drift-adaptive online learners.
- Bridging three literatures: The work connects survival analysis, longitudinal/joint modelling and data drift adaptation, with the stated claim that the framework extends beyond mortgage default to other operational risk and reliability problems with evolving behaviour and drifting distributions.
Main Findings
- Drift matters and is explicitly simulated: Concept drift (sudden, incremental, recurring) and label drift are injected simultaneously and independently into the Freddie Mac data, applied to the full time-indexed monthly panel before cross-validation folds are constructed, so every loan in every fold carries a drifted trajectory. This design tests whether a model trained on a representative portion of simulated data generalises to unseen loans under the same non-stationarity.
- Proposed model wins across drift scenarios (per abstract): The landmark-based joint model "consistently outperforms classical survival models, tree-based drift-adaptive learners and gradient boosting methods in terms of discrimination and calibration across all drift scenarios."
- Behavioural marker construction: The key longitudinal signal is BD_pct(t), the percentage deviation of the current actual unpaid balance from the scheduled balance implied by standard amortisation. Positive values indicate the borrower is behind schedule, negative values indicate ahead of schedule. The authors argue this construction is economically interpretable and scale-free across loans with different original balances and terms.
- Latent state summarised by per-loan regression: Each loan's BD_pct trajectory is summarised by a two-parameter linear fit in normalised age (t/N_i), giving an intercept b0i (early deviation / level) and a slope b1i (deterioration or improvement as the loan seasons), estimated by closed-form OLS with a small ridge term (lambda) for stability.
- Drift-severity diagnostics: For numeric variables, drift level is measured as per-loan median month-over-month absolute change normalised by the variable's global IQR, with thresholds None (0, 0.1), Slight (0.1, 0.3), Moderate (0.3, 0.7), Severe above 0.7. For categorical variables, per-loan state-change rate is used with thresholds None (0, 0), Slight (0, 0.1), Moderate (0.1, 0.3), Severe (0.3, 1.0). The paper reports these distributions for 2020 as the illustrative year.
- Threshold sensitivity: Because the target is defined as CurLoanDel not equal to 0 (an early-delinquency definition rather than the Basel-style 90-days-past-due criterion), the authors ran a sensitivity analysis using CurLoanDel greater than or equal to 3 for the 2020 drift scenarios. They report that the 90-days-plus definition significantly reduces positive cases and depresses F1 for the hazard-based models even when AUC remains moderate, confirming that the two thresholds represent different operational tasks.
- Numerical results are not reported in the provided content: The specific AUC, Brier score, F1 or other metric values, and the exact ranking of competing models, appear in Section 5 of the paper, which is not included in the truncated content supplied.
Methodology in Plain English
The behavioural marker. Standard amortisation mathematics tells you what a loan balance should be at any month given the interest rate, term and original balance. The paper compares that scheduled balance with the balance the servicer actually reports. The resulting percentage gap, BD_pct, says whether the borrower is ahead of or behind the expected repayment path. For each loan, a simple straight-line fit of BD_pct against normalised loan age produces two numbers: where the borrower started relative to schedule, and whether they are drifting further behind or catching up over time.
Landmarking. Instead of training one model over the whole loan history, the authors pick "landmark" months. At each landmark L, they take all loans still alive, gather everything known up to L (static characteristics plus the behavioural marker), and ask whether a default occurs in the next H months. This converts one survival problem into a sequence of aligned prediction problems, each of which only uses information available at that point in time.
Discrete-time hazard model. The probability of default within the horizon is modelled with a logistic regression on the covariates and the behavioural marker. The coefficient on the marker measures how strongly diverging from the scheduled repayment path feeds back into default risk.
Landmark one-hot encoding (LM). Different landmark months have different average risk because of seasoning, survivorship and macro conditions. Adding one-hot indicators for the landmark month gives each month its own baseline intercept, while keeping the covariate effects shared. The paper describes how to handle new observations at intermediate months (assign to nearest landmark or interval) or beyond the training landmark range (retain the final category as conservative extrapolation, with periodic refitting recommended).
Isotonic calibration (ISO). Class imbalance and drift can distort predicted probabilities even when ranking is fine. Isotonic regression fits a non-decreasing mapping from raw predicted probabilities to observed outcomes by minimising squared error, which preserves rank ordering (so typically preserves AUC) while improving calibration metrics such as the Brier score.
Data and drift simulation. The study uses Freddie Mac's public Single Family Loan-Level Dataset (updated 30 June 2025), sampling 50,000 loans per year for computational tractability, with experiments on drifted datasets from 2000, 2010, 2020 and 2021. Origination data provides static covariates (CreditScore, Occupancy, DTI, OrigUPB, OrigLTV, OrigInterestRate, LoanPurpose, OrigLoanTerm, NumBorrowers); monthly performance data provides time-varying variables (CurAct_UPB, CurLoanDel, LoanAge, ZeroBalCode, CurIntRate, CNIB_UPB, ELTV, ASSISTANCE_CODE). Only CurAct_UPB, CurIntRate, ELTV and CurLoanDel are perturbed.
Drift schedules. Sudden drift: a one-time break at one-third of the observation window (t_s = floor(T/3)), shifting CurIntRate by +1, scaling ELTV by 1.2, and reducing the first post-break CurAct_UPB observation per loan by 5 percent. Incremental drift: a linear ramp from one-third to two-thirds of the window (t_s = floor(T/3) to t_e = floor(2T/3)), with CurIntRate plus 1.5·tau(t), ELTV times (1 + 0.15·tau(t)), and CurAct_UPB times (1 − 0.09·tau(t)), where tau(t) rises from 0 to 1. Recurring drift: a 12-month sinusoidal cycle, with CurIntRate plus 0.5·sin(2πt/12), ELTV times (1 + 0.05·sin(2πt/12 + π/6)), and CurAct_UPB times (1 − 0.02(0.5 + 0.5·sin(2πt/12 + π/3))). Label drift flips labels probabilistically per month: sudden shifts prevalence from 0.025 to 0.10 at t_s; incremental ramps prevalence linearly from 0.025 to 0.12 over the window; recurring oscillates as 0.06 + 0.035·sin(·). The illustrative schedule figure spans the first 60 months with t_s = 20 and t_e = 40.
Preprocessing. Loans that cannot be linked by LoanSeqNum across the two datasets are removed. Invalid values (negative balances, negative loan terms, non-positive original balances, loan terms above 1000 months, and interest-rate or LTV values outside economically meaningful ranges) are coded as missing rather than winsorised, because the authors consider extreme values more likely to indicate data quality problems than genuine borrower behaviour. Prepayment is not modelled as a competing risk.
Why This Matters
Impact on research. The paper pushes credit risk survival analysis from a stationary-offline framing toward an explicitly non-stationary one. It argues that robustness under drift requires an integrated strategy covering longitudinal borrower behaviour, time-to-event dynamics, and temporal heterogeneity in the risk distribution simultaneously — rather than treating these as separate problems. It also provides a benchmark protocol (simulated sudden, incremental and recurring concept drift plus label drift, applied before fold construction) that other researchers can reuse. The authors position joint modelling and landmarking as the two principal strategies for dynamic survival prediction with longitudinal data and adopt landmarking here, while citing work that systematically compares the trade-offs.
Real-world applications (as described by the authors):
- Portfolio monitoring: Regular updating of default probabilities as new monthly servicing information arrives, supporting ongoing surveillance of a mortgage book.
- Early warning systems: The early-delinquency target definition (any non-zero delinquency) is chosen so the model detects deterioration early enough for intervention rather than only flagging conventional Basel-style defaults.
- Stress testing: The drift-aware design is intended to hold up when economic, regulatory or portfolio-composition conditions change materially.
- Operational risk and reliability more broadly: The authors state the framework applies to other problems characterised by evolving behaviour and drifting data distributions, not just mortgage default.
Industry relevance. The approach is described as applicable to routinely collected servicing data, and the behavioural marker is presented as improving forward portability by capturing repayment behaviour in an economically interpretable way rather than relying only on fixed historical patterns. The paper is candid about the operational cost: because of the time-specific adjustments and probability calibration, the model should be recalibrated regularly, especially after major economic, regulatory or portfolio-composition changes, and refitted when portfolios move into materially new age ranges or macroeconomic regimes.
Future Directions
- Replace the lightweight per-loan linear trajectory with a full random-effects or latent-process joint model. The paper explicitly describes its two-parameter OLS fit as "a computationally lighter random-effects proxy," leaving open whether richer latent longitudinal structures improve prediction under drift.
- Extend calibration and drift handling to competing risks. Prepayment is deliberately not modelled as a competing risk in this analysis, which is a stated scope limitation; adding competing-risk structure is a natural extension for real portfolios where prepayment and default interact.
- Move from simulated drift to observed regime changes. The drift regimes here are parametric and controlled by design; validating the framework against naturally occurring macroeconomic shifts and policy interventions is the next empirical step.
- Establish the right operational default definition. The sensitivity analysis using CurLoanDel greater than or equal to 3 versus CurLoanDel not equal to 0 shows the two thresholds correspond to different operational tasks; deciding which threshold a deployment should target, and how that choice interacts with drift adaptation, remains an open question.
Target Audience
Quantitative risk modellers and model-validation teams in retail banking and mortgage lending; credit risk researchers working on survival analysis, joint modelling or dynamic prediction; machine learning researchers interested in concept drift in imbalanced, time-to-event settings; and regulators or auditors concerned with model stability, probability calibration and recalibration policy under changing economic conditions. Readers wanting the concrete benchmark numbers should go to Section 5 of the full paper, which is not contained in the content summarised here.
Authors’ abstract
Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk. Unlike most existing methods that implicitly assume a stationary data-generating process, in practise, mortgage portfolios are exposed to various forms of data drift caused by changing borrower behaviour, macroeconomic conditions, policy regimes and so on. This study investigates the impact of data drift on survival-based credit risk models and proposes a dynamic joint modelling framework to improve robustness under non-stationary environments. The proposed model integrates a longitudinal behavioural marker derived from balance dynamics with a discrete-time hazard formulation, combined with landmark one-hot encoding and isotonic calibration. Three types of data drift (sudden, incremental and recurring) are simulated and analysed on mortgage loan datasets from Freddie Mac. Experiments and corresponding evidence show that the proposed landmark-based joint model consistently outperforms classical survival models, tree-based drift-adaptive learners and gradient boosting methods in terms of discrimination and calibration across all drift scenarios, which confirms the superiority of our model design.