Skip to content
AI.info

AI.info

ML data engineering

The evidence platform underneath every model: prediction time, ownership, contracts, and the leaks that look like accuracy.

  1. ML Data Engineering as an Evidence System
  2. Prediction Time, Unit of Analysis, and Outcome Horizon
  3. Source Discovery, Ownership, and Data Authority
  4. Semantic Contracts and Business Invariants
  5. Learning Event Instrumentation and Taxonomy
  6. Clocks: Event, Ingestion, Availability, and Label Time
  7. Batch, Change Data Capture, and Stream Ingestion
  8. Windows, Watermarks, and Late Data
  9. Storage Layers, File Formats, and Table Layout
  10. Table Snapshots, Compaction, and Retention
  11. Schema Evolution, Migrations, and Historical Backfills
  12. Identifiers, Keys, and Entity Resolution
  13. Joins, Cardinality, Deduplication, and Reconciliation
  14. Training Dataset Construction and Release
  15. Point-in-Time Correctness and Historical Retrieval
  16. Labels, Ground Truth, and Annotation Pipelines
  17. Label Delay, Censoring, and Proxy Targets
  18. Data Profiling and Statistical Baselines
  19. Data Quality Rules, Constraints, and Release Gates
  20. Missingness, Outliers, Corruption, and Repair
  21. Feature Transformations and Stateful Pipelines
  22. Data Engineering for Text, Images, Audio, and Sensors
  23. Categorical, Sparse, and High-Cardinality Data
  24. Split Design, Sampling, and Leakage Prevention
  25. Imbalanced Data, Rare Events, and Coverage
  26. Time-Series and Sequential Data Engineering
  27. Graph Data Engineering and Dynamic Relations
  28. Data Augmentation, Synthetic Data, and Weak Supervision
  29. Feature Stores and Historical Retrieval
  30. Online Serving Data, Caching, and Fallbacks
  31. Dataset Versioning, Lineage, Provenance, and Reproducibility
  32. Orchestration, Idempotency, Backfills, and Recovery
  33. Data Validation and Release Gates
  34. Data Observability, SLIs, and Incident Response
  35. Drift, Freshness, Refresh, and Dataset Retirement
  36. Privacy, Consent, Retention, and Deletion
  37. Data Security, Poisoning, and Supply-Chain Risk
  38. Dataset Documentation, Governance, and Use Review
  39. Cost, Performance, and Data Platform Design
  40. Capstone: Design and Defend an ML Evidence Platform