Skip to content
AI.info

AI.info

MLOps

Production as a control system — contracts, risk tiers, ownership, and what happens at three in the morning.

  1. MLOps as a Control System
  2. Production Contracts, Risk Tiers, and Ownership
  3. ML Workload Architecture: Batch, Online, Streaming, and Edge
  4. Reproducible Environments, Configuration, and Secrets
  5. Versioning the ML Asset Graph
  6. Lineage, Provenance, and Release Evidence
  7. Experiment Tracking and Decision Records
  8. Registries, Approval, and Release Candidates
  9. Pipeline Design and Orchestration
  10. Idempotency, Caching, Backfills, and Recovery
  11. Testing Machine Learning Systems
  12. Model Packaging and Runtime Contracts
  13. Batch Inference Systems
  14. Online Inference Services
  15. Streaming and Event-Driven Inference
  16. Edge and On-Device MLOps
  17. Continuous Integration for ML Assets
  18. Continuous Delivery and Progressive Release
  19. Continuous Training and Retraining Policy
  20. Observability for ML Systems
  21. Service-Level Objectives and Error Budgets for ML
  22. Production Data Quality and Training–Serving Skew
  23. Model Quality Monitoring, Drift, and Delayed Labels
  24. Feedback, Labeling, and Human Review Loops
  25. ML Incident Response, Rollback, and Disaster Recovery
  26. ML Security and Supply-Chain Integrity
  27. Privacy, Governance, and Audit Operations
  28. Production Explainability, Documentation, and Decision Evidence
  29. Capacity, Cost, Performance, and Sustainable ML Operations
  30. ML Platform Engineering, Self-Service, and Multi-Tenancy
  31. Operating Foundation and Generative AI Systems
  32. Model Retirement, Decommissioning, and Evidence Retention
  33. MLOps Capstone: Design and Defend a Production System