Skip to content
AI.info

MLOps

MLOps Capstone: Design and Defend a Production System

Integrate architecture, evidence, delivery, observability, incident response, governance, economics, and retirement into one defended system design.

By the end you can

The scenario: wildfire-risk inspection for an electric utility

An electric utility wants to prioritize inspections of poles and vegetation using aerial images, weather, maintenance history, and sensor data. The ranked list will influence field crews, emergency preparation, and maintenance budgets.

The inspection regime that list feeds into is not an informal habit, and it is not yours to define. In California the rules it can be judged against carry numbers. On 26 November 2019 the Safety and Enforcement Division of the California Public Utilities Commission moved to widen its investigation to cover the 2018 Camp Fire, and attached a table of the violations it had found: General Order 95 Rules 18, 31.1, 31.2 and 44.3, General Order 165 Section IV, Resolution E-4184, and Public Utilities Code § 451. One of them was the failure to conduct detailed climbing inspections when internal triggers were evident.

What an inspection coverage gap looks like in that regime is also on the record. The Camp Fire ignition point was Tower 27/222 of PG&E's Caribou-Palermo 115kV line, whose transposition hardware had been in service since 1921. The Butte County District Attorney, Michael L. Ramsey, worked through every inspection and patrol record he could obtain back to 2001. His 92-page public report of 16 June 2020 found nothing there at all: “There is no record of any climbing inspections, detailed ground inspections above 10’ or aerial inspections conducted on the Caribou-Big Bend section of the transmission line.” The same report documents that PG&E had “little or no information about the 97-year-old conductor”, and that a 1995 ES Guideline had eliminated routine climbing inspections.

That is the ground your ranked list lands on: a numbered standard, a documented coverage failure, and a regulator willing to enumerate which rule each gap broke. Labels are delayed and incomplete. Remote regions have weak connectivity. Wildfire season creates sharp demand peaks, and the company must explain its decisions after incidents. So the capstone is to design the operating system around the model, not to choose the most impressive architecture.

Case

The fire downstream of an inspection list

Wildfire-risk inspection is not a hypothetical exercise. The Camp Fire ignited on 8 November 2018 in Butte County, California. NIST's case study of it records “over 18,000 destroyed structures, 700 damaged structures, and 85 fatalities”. CAL FIRE's own incident record gives 153,336 acres burned, 18,804 structures destroyed and 85 civilian fatalities. It lists the cause as power lines.

On 15 May 2019 CAL FIRE announced that the fire “was caused by electrical transmission lines owned and operated by Pacific Gas and Electricity (PG&E) located in the Pulga area”. PG&E put it in its own words the same day, in a news release filed with the SEC: “CAL FIRE announced today that it has determined that PG&E electrical transmission lines near Pulga were a cause of the Camp Fire.”

A ranked inspection list sits upstream of decisions with that consequence. The tower at the top of it was one the records show nobody had climbed.

Example

Evidence packets the review board will demand

The capstone should produce artifacts, not only an architecture diagram. Each packet below exists because some real review — a district attorney's investigation, a regulator's violation table, a court's footnote — has already asked for exactly that document and found it missing somewhere.

  • Production contract: Population, decision deadline, ranked-list semantics, exclusions, capacity, fallback, and prohibited automation. Write down which of those General Order obligations the list is meant to help discharge, and which it explicitly does not.
  • Release bill of materials: Data, features, model, runtime, policy, evaluation, approval, and rollback identities. Make them specific enough to tell whether all eight servers are running the same build. That is the check Knight Capital did not have.
  • Evaluation plan: Temporal and geographic splits, rare-event metrics, calibration, crew-capacity thresholds, and independent field audits. The Epic Sepsis Model shipped with vendor AUCs of 0.76-0.83 and measured 0.63 in an external validation. Assume your own numbers will be re-measured by someone who did not build them.
  • Operations plan: SLOs, feature freshness, traceability, alert runbooks, incident severity, reconciliation, and disaster recovery. The Butte County District Attorney went looking for inspection records back to 2001 and found none for the relevant section. Your traceability has to survive that kind of retrospective request.
  • Lifecycle plan: Retraining triggers, provider and sensor change, security review, retirement, evidence retention, and deletion — including the post-market monitoring and corrective action that Regulation (EU) 2024/1689 attaches to high-risk systems after they are placed on the market.

Visual

The required architecture

The proposal should connect each layer to a requirement and recovery path: evidence and learning, release control, execution, operations, and governance and lifecycle.

The governance layer is the one teams most often draw as an empty box. It is also the one with a statute already written into it. The AI Act, Regulation (EU) 2024/1689, was published in the Official Journal on 12 July 2024 and entered into force on 1 August 2024. The European Parliamentary Research Service gives the citation as “Regulation (EU) 2024/1689, OJ L, 2024/1689, 12.7.2024”. Its Annex III, the list Article 6(2) points to, names at point 2: “Critical infrastructure: AI systems intended to be used as safety components in the management and operation of critical digital infrastructure, road traffic, or in the supply of water, gas, heating or electricity.”

A model that ranks inspections for an electricity supplier is therefore not in a governance vacuum, and “governance” is not a box to be filled in later. The classification carries obligations on risk management, data governance, human oversight and post-market monitoring, before and after the system is placed on the market. Each of those obligations should terminate in a named artifact somewhere in your architecture graph, with an owner. Not in a colored rectangle.

FigureLayers · 5 layers
  1. 01

    Evidence and learning

    Source contracts, point-in-time datasets, labels, evals, and retraining policy.

  2. 02

    Release control

    Version graph, registry, gates, approval, canary, and rollback.

  3. 03

    Execution

    Batch planning, online review, edge fallback, and workload SLOs.

  4. 04

    Operations

    Observability, data quality, delayed labels, incident response, and capacity.

  5. 05

    Governance and lifecycle

    Security, privacy, explanation, appeal, evidence retention, and retirement.

Comparison

Three defensible deployment positions

The project is not required to end with full automation. Three positions can be defended: decision support only, constrained automation within a narrow eligible scope, and no deployment yet.

The gap between the first two and the third is measurable, and one agency measured it on its own program. On 11 August 2017 Michigan's unemployment agency published a review of 62,784 unappealed fraud-penalty cases from October 2013 to August 2015. Where the MiDAS computer program had resolved the case by itself, the finding was reversed 85 percent of the time: “Of those cases, 40,195 were originally resolved by way of computer program based on available information. As part of the review, 85 percent of these original fraud findings were reversed.” The other 22,589 cases were initiated by the computer and then reviewed by an investigator. Those were reversed at 44 percent. Same program, same period, same population — 85 percent against 44 percent, and refunds of more than $20.8 million.

The Michigan Supreme Court put the human cost of the automated branch on the record in Bauserman v Unemployment Insurance Agency, decided 26 July 2022. Roughly 40,000 people were wrongly accused, and footnote 5 notes that “a study conducted by the Agency concluded that, during this same period, approximately 93% of the automated system's fraud determinations were incorrect”. A reviewer in the loop is not a rhetorical concession in a design memo. In the only place anyone measured it here, it halved the error the system shipped.

FigureComparison · 3 columns

Decision support only

Provide ranked evidence to qualified planners who retain authority.

  • Lower automation risk
  • Needs reviewer capacity and interface design
  • Can collect better feedback
  • Appropriate when outcome evidence is incomplete

Constrained automation

Automate narrow low-risk actions under hard rules and audit.

  • Requires explicit eligible scope
  • Strong fallback and reconciliation
  • Continuous outcome review
  • Useful when evidence is mature in selected regions

No deployment yet

Keep baseline and improve data, workflow, or evaluation first.

  • Prevents premature complexity
  • Can run shadow evaluation
  • Focuses investment on evidence gaps
  • Correct when benefits are not defensible

The design must survive incomplete evidence

Confirmed failure and fire outcomes are rare, while inspection findings are selected by previous policy. Weather, sensors, contractors, and image sources change. The system therefore needs proxy monitoring, independent audits, mature outcome cohorts, and explicit limits on what the model can claim.

What an independent audit does to a deployed model's own numbers is documented. The Epic Sepsis Model was in use at hundreds of US hospitals when Wong and colleagues validated it from outside, in JAMA Internal Medicine in 2021. Across 38,455 hospitalizations of 27,697 patients at Michigan Medicine, the hospitalization-level AUC was 0.63 (95% CI, 0.62-0.64). The vendor's own documentation said 0.76-0.83. The operational shape of that gap is in the abstract: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” Their conclusion was that “the ESM has poor discrimination and calibration in predicting the onset of sepsis”.

One audit can be argued with. A second one, at different institutions, is harder. Ostermayer and colleagues ran the same model against two county emergency departments and reported in JAMIA Open in 2024 that “Within a 6-hour time window for sepsis, the ESPMv1 had a sensitivity of 14.7%, specificity of 95.3%, positive predictive value of 7.6%, and negative predictive value of 97.7%.” Two independent groups, one already-deployed model, and figures the vendor's documentation did not lead anyone to expect.

A valid conclusion may therefore be to deploy only as decision support, restrict certain regions, or retain a rules-based baseline until evidence improves. The evidence gap is not resolved by shipping and watching a dashboard. It is resolved by someone outside the build measuring the thing.

Analogy

The design review is an aviation safety case

This capstone is a safety case for a new aircraft route. The team must show the vehicle, route, weather limits, crew procedures, maintenance, incident response, and evidence supporting operation. A strong engine test alone does not authorize the route.

The analogy carries a warning with it, because formal safety cases fail in a specific way. RAF Nimrod MR2 XV230 was lost in Afghanistan on 2 September 2006 with its 14 crew. The independent review that followed, published on 28 October 2009, examined the aircraft's formal Safety Case. It found 40% of the hazards left “Open” and a further 30% “Unclassified”, and the whole thing reduced to “essentially a paperwork and 'tick-box' exercise”. Its author, Charles Haddon-Cave QC, did not hedge the verdict: “Unfortunately, the Nimrod Safety Case was a lamentable job from start to finish. It was riddled with errors. It missed the key dangers. Its production is a story of incompetence, complacency, and cynicism.”

A safety case with seven in ten hazards open or unclassified still looked like a safety case. That is the failure mode a capstone document is most exposed to. Count your unresolved hazards. Do not describe your process.

An approved route also stays approved because the physics under it holds still. The population a model scores keeps moving, and the evidence about that population ages with it. That is why this safety case ends in monitoring, change control, and named retirement triggers, rather than in one static approval.

The capstone defends a controlled operating claim, not the abstract intelligence of a model.

Key idea

The capstone fails if it hides uncertainty behind architecture

A polished diagram cannot compensate for selected labels, weak geographic coverage, undefined inspection capacity, or missing rollback. Every major component should answer a stated requirement and name the evidence needed to trust it.

The cost of drawing around uncertainty is paid by people who never saw the diagram. The Michigan Supreme Court described MiDAS as an “automated fraud-detection system” that “can result in recipients being disqualified from benefits and subjected to penalties and criminal prosecution, all without notice or an opportunity to be heard”. Nothing in that sentence is about model quality. It is about which authority the system was handed, and which safeguards were removed alongside it. Those are decisions made in a design document, by someone with a diagram.

Do not invent guarantees. State unresolved questions, assumptions, excluded uses, and conditions that would stop or reverse deployment. Then count them, rather than characterize them.

A defensible system design exposes uncertainty and authority instead of drawing around them.

Steps

Build the capstone in seven passes

Each pass produces an artifact that can be reviewed independently: the product and risk contract, the asset and architecture graph, the evidence and experiment plan, release and recovery, operations and incident response, security and governance, and the decision memo.

Pass 4 is the one teams treat as plumbing, and it has the sharpest documented failure. Knight Capital Americas rolled new RLP code out to eight servers. One server did not get it and kept running discontinued Power Peg code. The SEC's order of 16 October 2013 sets out how: “During the deployment of the new code, however, one of Knight’s technicians did not copy the new code to one of the eight SMARS computer servers. Knight did not have a second technician review this deployment and no one at Knight realized that the Power Peg code had not been removed from the eighth server, nor the new RLP code added. Knight had no written procedures that required such a review.”

On 1 August 2012 that eighth server turned 212 incoming parent orders into 4 million executions in 154 stocks, for more than 397 million shares, in about 45 minutes. The loss was $460 million. Knight paid a $12,000,000 civil money penalty. Carol Clark, writing for the Federal Reserve Bank of Chicago, named the control that was absent: “Trading firms should have access controls defining who can develop, test, modify, and place an algorithm into production, as well as quality assurance procedures for the code and development processes.” A canary and a rollback plan in pass 4 are not ceremony. Here the entire failure was one machine out of eight that nobody was required to check.

FigureProcess · 7 steps
  1. 1. Product and risk contract

    Define decisions, users, consequences, populations, and non-negotiable controls.

  2. 2. Asset and architecture graph

    Map data, pipelines, models, workloads, policies, dependencies, and owners.

  3. 3. Evidence and experiment plan

    Specify splits, baselines, labels, slices, uncertainty, and acceptance criteria.

  4. 4. Release and recovery

    Design CI, candidate bundle, approval, canary, rollback, and reconciliation.

  5. 5. Operations and incident response

    Set SLOs, telemetry, quality monitoring, runbooks, capacity, and disaster recovery.

  6. 6. Security and governance

    Threat-model authority, privacy, explanation, appeals, audit, and retention.

  7. 7. Decision memo

    Recommend support, constrained automation, or no deployment with reversal conditions.

The final deliverable is a decision, not a technology stack

The decision memo should compare alternatives, identify the safest useful scope, and state why the chosen complexity is justified. It should carry expected benefit, residual risk, evidence gaps, owners, stop conditions, and retirement criteria.

Approval is not the end of the obligation, and in this domain that is now written into law rather than into good practice. For high-risk systems under Regulation (EU) 2024/1689, the European Parliamentary Research Service records that once a system is placed on the market “providers must implement post-market monitoring and take corrective actions if necessary”. A memo that ends at launch has answered a question nobody with authority asked.

A recommendation not to deploy can earn full credit when it follows from the evidence. Michigan's own review of MiDAS, the two external validations of the Epic Sepsis Model, and the Nimrod Safety Case with 40% of its hazards still open are three different institutions discovering, after the fact, what a pre-deployment memo could have written down in advance. MLOps maturity includes the ability to prevent an unjustified system from becoming operational.

Case

Eighty-four counts, in a Butte County courtroom

Accountability for wildfire risk has been criminal, not merely reputational. On 16 June 2020 — the same day the District Attorney released his public report on the inspection records — PG&E stated that it had “pleaded guilty to 84 counts of involuntary manslaughter and one count of unlawfully starting a fire”. The 84 felony counts of involuntary manslaughter and the single felony count of unlawfully starting a fire were entered before Judge Michael Deems in Butte County Superior Court.

That is the room a decision memo is eventually read in, next to the maintenance records, the inspection history, and the rule numbers those records were supposed to satisfy. It is a good reason to write down the conditions for not operating the system, before anyone asks.

Key takeaways