Memory hook: Split before learning transformations; evaluate before deployment; monitor after deployment.
Must remember
- Start with a measurable business problem, an acceptable error cost and a baseline. Collect authorised, representative data; clean missing/invalid values; transform features; split training, validation and test sets; train; evaluate; deploy; monitor and retrain when justified.
- Fit scalers, imputers and feature selection on training data, then apply the learned transformation to validation/test data. Duplicate entities or future events across splits can create data leakage and unrealistic scores.
- Overfitting means learning training-specific patterns that generalise poorly; regularisation, representative data and simpler models can help. Underfitting means the model cannot capture useful patterns; improve features or capacity. Hyperparameters control training choices; model parameters are learned.
- Precision = TP/(TP+FP): of predicted positives, how many are correct? Recall = TP/(TP+FN): of actual positives, how many were found? F1 combines precision and recall. Accuracy can hide failure on a rare class. For regression, error metrics must match the business penalty.
- Batch inference handles offline collections; real-time endpoints serve interactive calls; asynchronous inference accepts queued work with later results; serverless inference reduces endpoint-management effort for supported patterns. Match latency, payload, traffic and cost.
- MLOps versions data, code and model artifacts; tracks experiments; automates repeatable pipelines; approves releases; monitors drift and operating health. SageMaker AI supports these lifecycle stages. A technically better score is insufficient if cost per user rises beyond the value of the improvement.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Missing a disease is especially costly | Prioritise recall while measuring the resulting false positives. |
| Many alerts are wrong | Investigate precision and the decision threshold. |
| Millions of records need results tomorrow | Evaluate batch inference. |
Traps
- The test set is not a tuning dataset.
- Data drift does not prove accuracy declined; collect outcome evidence.
- A model needs operational and business metrics as well as statistical scores.
Active recall
1. TP=80 and FP=20: what is precision?
80/(80+20)=80%.
2. TP=80 and FN=20: what is recall?
80/(80+20)=80%.
3. Why is 99% accuracy weak evidence when only 1% of events are fraud?
Always predicting legitimate already achieves 99%; inspect minority-class performance.
4. Training improves while validation worsens. What is likely?
Overfitting; investigate leakage, regularisation and model complexity.
5. Why retain a model version after replacing it?
For reproducibility, audit and a tested rollback path.