Skip to content
AI.info

Evaluation

Anomaly Detection and Alerting Evaluation

Evaluate anomaly detectors through alert yield, event-level coverage, detection delay, false-alert burden, and realistic temporal backtests.

By the end you can

Most normal points make accuracy look excellent

If one in 100,000 events is harmful, a system predicting “normal” for everything reaches 99.999% accuracy. It detects nothing.

Anomaly evaluation asks other questions. How many of the rare events were found? How many alerts did that cost? How long did detection take, and who has to work through the queue?

Every one of those has a published number attached to it later in this lesson. 88.8% of adjudicated arrhythmia alarms false in one intensive care unit. 67% of sepsis patients never flagged by a deployed model. 275 alarms in the last 11 minutes before a refinery explosion. Accuracy looked respectable throughout. So, in one case, did a specificity of 95.3%.

A detector exists to find rare structure, so majority-class accuracy is nearly irrelevant.

Case

A random anomaly score that scored like the state of the art

Many papers apply point adjustment before computing F1. What that does to a leaderboard is not subtle. Under the protocol, Kim and colleagues write, “even a random anomaly score can easily turn into a state-of-the-art TAD method”, and comparisons made afterwards “can lead to misguided rankings”.

The finding does not even depend on the protocol being applied. Their next line: “an untrained model obtains comparable detection performance to the existing methods even when PA is forbidden”. The paper appeared at AAAI in 2022.

A metric that a random score and an untrained model can both saturate is measuring the protocol, not the detector.

Visual

The evaluation unit changes with anomaly type

Point-level scoring is not always the right contract. One 31-day stretch of intensive care monitoring, counted below, is either 2,558,760 events or a few hundred clinical episodes. The raw data is identical. Only the unit of analysis differs.

The two counts support opposite conclusions about the same detector. So the unit has to be chosen and stated before any metric is computed.

FigureHierarchy · 4 levels
  • Point anomaly

    One observation is unusual relative to a reference distribution.

    • Contextual anomaly

      A value is unusual only given time, location, user, or operating state.

      • Collective anomaly

        A sequence or group is suspicious even if individual points look ordinary.

        • Incident event

          Many anomalous observations belong to one operational episode.

Comparison

Alert metrics versus event metrics

An incident can generate hundreds of alerts, and the four metrics below routinely move in opposite directions on one detector. Epic's sepsis model has been externally validated twice. The two studies fill the table with real values.

The first validation covered 38,455 hospitalizations and put the model's AUC at 0.63. It “did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%)”, at a number needed to evaluate of 8. Wong and colleagues reported that in JAMA Internal Medicine in 2021.

An independent cohort of 145,885 emergency department encounters supplies the rest of the table. “Within a 6-hour time window for sepsis, the ESPMv1 had a sensitivity of 14.7%, specificity of 95.3%, positive predictive value of 7.6%, and negative predictive value of 97.7%.” On timing, Ostermayer and colleagues add: “Providers were alerted with a median lead time of 0 minutes (80% CI, −6 hours and 42 minutes to 12 hours and 0 minutes).” That one ran in JAMIA Open in 2024.

Read against the four columns, one deployed detector scores alert precision of 7.6%. Its event recall leaves 1,709 sepsis patients unflagged. Its median detection delay is zero minutes — the alert arrives as the clinician already knows. Its burden falls on 18% of everyone admitted. No single one of those numbers would have exposed the others.

FigureComparison · 4 columns

Alert precision

Fraction of alerts linked to true incidents.

  • Measures investigator yield
  • Sensitive to deduplication
  • Depends on label completeness
  • Does not measure event coverage

Event recall

Fraction of true incidents detected at least once.

  • Matches incident coverage
  • Needs event definitions
  • Can ignore alert flood
  • Should include severity

Detection delay

Time from incident onset to first actionable alert.

  • Critical for intervention
  • Needs onset labels
  • Can trade against false alerts
  • Often right-censored

Alerts per unit time

Operational burden per hour, device, user, or site.

  • Connects to capacity
  • Easy to monitor
  • Does not measure usefulness
  • May vary with traffic

Key idea

Incomplete labels distort false-positive counts

Investigators usually examine only selected alerts, so unreviewed cases lack outcomes, and some incidents are never discovered, while known-event datasets overrepresent severe or easy cases.

Treat missing labels as missing, not automatically normal. Use delayed adjudication, sampled review, synthetic injections, weak labels, and prospective incident tracking with stated limitations.

Target is the documented case. A Senate Commerce Committee majority staff report in 2014 found that “Target appears to have failed to respond to multiple automated warnings from the company's anti-intrusion software that the attackers were installing malware on Target's system.” The report is a synthesis of the public record, not an independent forensic investigation.

The alerts themselves fired on schedule. Target's FireEye intrusion detection system triggered an urgent alert at each installation of the data-exfiltration malware — first installed 30 November 2013, updated just before midnight on 2 December and just after midnight on 3 December, per Dell SecureWorks. What happened next takes the report one sentence: “However, Target’s security team neither reacted to the alarms nor allowed the FireEye software to automatically delete the malware in question.” Target officials testified that they were unaware of the breach until the Department of Justice contacted them on 12 December 2013. An independent technical reconstruction in 2017 reached the same line: “Unfortunately, multiple malware alerts were ignored.”

Every metric computed from that alert log would have been wrong in a specific direction. The detector's measured precision was depressed by alerts nobody adjudicated. Its measured detection delay was twelve days. Its actual lead time was zero.

Unreviewed does not mean non-anomalous.

Example

Two counted alert floods

Point-level scoring has been measured against operational reality in two settings where somebody kept the tally: a month of intensive-care monitoring, and the eleven minutes before an explosion.

  • Volume: 461 consecutive adult intensive care patients at UCSF were monitored for a month, and every alarm was counted. “A total of 2,558,760 unique alarms occurred in the 31-day study period: arrhythmia, 1,154,201; parameter, 612,927; technical, 791,632.” Drew and colleagues published the tally in PLoS ONE in 2014.
  • Burden: 381,560 of those alarms were audible. That is 187 per bed per day. The Agency for Healthcare Research and Quality restates the same figure from its own citation in its primer on alert fatigue: “A 2014 study found that the physiologic monitors in an academic hospital's 66 adult intensive care unit beds generated more than 2 million alerts in one month, translating to 187 warnings per patient per day.”
  • Yield: nurse scientists annotated 12,671 of the arrhythmia alarms one by one. 88.8% were false positives. That is the alert-precision column of the previous section, measured by hand.
  • Flood: an explosion and fires at the Texaco refinery at Milford Haven on 24 July 1994 injured 26 people and caused around £48 million of damage. The UK Health and Safety Executive recorded what the control room faced: “In the last 11 minutes before the explosion the two operators had to recognise, acknowledge and act on 275 alarms.” A 2021 European Process Safety Centre deck puts the same figure per second: “In the last 11 minutes before the explosion the two operators had to deal with 275 alarms. Almost an alarm each 2.4 seconds.”
  • Grouping: the Institution of Chemical Engineers' incident summary lists among critical factors that the “Control board operator was overwhelmed by alarm flood in an emergency situation”, and among root causes “Inadequate warning systems (too many alarms, poorly prioritised)”. 275 true positives. One event nobody could act on.

Analogy

A smoke alarm in a building with many sensors

During one fire, every smoke sensor in the building sounds, and most of them sound repeatedly. Counting each sound as a separate detection exaggerates success and ignores whether the fire was identified quickly.

Hospitals do not have to imagine that building. Almost none of their alarms need a response: “It is estimated that between 85 and 99 percent of alarm signals do not require clinical intervention, such as when alarm conditions are set too tight; default settings are not adjusted for the individual patient or for the patient population; ECG electrodes have dried out; or sensors are mispositioned.” That is the Joint Commission, in a 2013 sentinel event alert on medical device alarm safety.

The same alert counts the harm. Between January 2009 and June 2012 the Joint Commission's Sentinel Event database logged 98 alarm-related events. 80 resulted in death, 13 in permanent loss of function, five in unexpected additional care or an extended stay.

A separate literature base gives a comparable figure. Reviewing alarm fatigue for the Agency for Healthcare Research and Quality, Woo and Bacon write: “Studies have shown that the percentage of false alarms can range from 72 percent to 99 percent.”

So the repeated beep is not a rhetorical flourish. A non-actionable alert rate is a measured property of a deployed detector, and in this instance a fatal one.

The brigade wants two things from that night. It wants to know that the fire was found, and it wants to know how long that took. Alerts have to be grouped into events before either question has an answer.

The operational unit is the incident, not every repeated beep.

Steps

Backtest incidents without looking into the future

Temporal leakage can make anomaly detectors appear clairvoyant. It also destroys the one quantity that decides whether an alert was worth raising: the interval between onset and the first actionable alert.

For the sepsis model already running in two emergency departments, that interval came out at a median of 0 minutes. The number only exists because the replay respected timestamps.

FigureProcess · 5 steps
  1. 1. Define event windows

    Mark onset, active period, recovery, severity, and exclusion rules.

  2. 2. Freeze history

    Fit normal profiles using only data available before each evaluation period.

  3. 3. Replay sequentially

    Generate scores and alerts in timestamp order with production state.

  4. 4. Group alerts

    Apply the same deduplication and suppression logic used operationally.

  5. 5. Score outcomes

    Measure event recall, delay, alert yield, burden, and slice behavior.

Thresholds should be expressed in capacity and incident terms

A threshold that produces 0.1% false positives can still overwhelm a system processing billions of events. Express burden per investigator, device, hour, or customer, and pair it with event coverage.

The process industries stopped recommending that and started publishing the rate. Following EEMUA Publication 191, the UK Health and Safety Executive states the contract: “Alarm rate targets: the long-term average alarm rate during normal operation should be no more than one every ten minutes; and no more than ten displayed in the first ten minutes following a major plant upset”.

The neighbouring standards agree on the shape. A 2022 review tabulates their limits side by side: IEC 62682 at no more than 2 per 10 minutes, ANSI/ISA-18.2 at fewer than 2 per 10 minutes, ISO 11064-5 at no more than 1 per 6 minutes, and Norwegian PSA YA-711 at one per 5 minutes. Not one of those five figures is a false-positive percentage. Every one of them is a rate per operator. That is the only form in which a threshold can be checked against the people who have to act on it.

Monitor drift in score distributions and alert mix, but remember that drift is not itself an incident label.

The benchmarks have been audited too. Most papers, Wu and Keogh observe, test on “a handful of popular benchmark datasets, created by Yahoo, Numenta, NASA, etc.”, and “The majority of the individual exemplars in these datasets suffer from one or more of four flaws”. Their verdict on the field's trajectory is that “much of the apparent progress in recent years may be illusionary”. Rather than stop there, they “introduce the UCR Time Series Anomaly Archive” as a replacement. The audit went up in 2020 and was published in IEEE Transactions on Knowledge and Data Engineering.

Rare-event evaluation becomes meaningful when rates are translated into workload and time.

Key takeaways