Memory hook: Validate before train; evaluate before promote.
Must remember
- An ML pipeline should include ingestion, validation, transformations, training, evaluation, registration and controlled deployment. Version component code, environments and inputs.
- Managed pipelines and Kubeflow-style components express dependencies and artifacts; Airflow coordinates broader workflows; Ray supports distributed computation patterns. Choose by orchestration and execution needs rather than treating them as interchangeable.
- Reuse the same preprocessing specification in training and serving. Data validation catches schema/distribution issues before they become misleading model updates.
- CI tests code/components; CD promotes approved artifacts; continuous training retrains on a defined trigger. A retraining trigger can be time, new data, drift or measured quality degradation.
- Require evaluation gates and lineage before promotion. A freshly trained model may be worse; use a champion/challenger comparison and a rollback plan.
- Use least-privileged pipeline identities, bounded retries, safe caching and artifact retention. Cache reuse is valid only when the relevant inputs and code are equivalent.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| New labelled data arrives weekly | A controlled retraining pipeline with validation and promotion gates. |
| Same preprocessing repeated inconsistently | A versioned reusable component shared by training and serving. |
Traps
- Retraining is not the same as automatically deploying every new model.
- A successful pipeline execution does not prove the model is better.
Active recall
1. What does continuous training automate?
Rebuilding models from defined inputs/triggers while preserving evaluation controls.
2. Why record artifact lineage?
To trace a deployed model to data, code and evaluation evidence.
3. When is cached component output unsafe?
When hidden dependencies or changed inputs are not represented in the cache key.
4. What should a failed data check do?
Stop or quarantine the affected workflow rather than promote an invalid model.
5. Why separate training and release permissions?
To prevent an experiment or compromised training job from silently replacing production.