certslothcertsloth
MLA-C02/Topic 01

AWS / Associate

The ML Lifecycle and MLOps

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Split before learning transformations; evaluate before deployment; monitor after deployment.

Must remember

  • Start with a measurable business problem, an acceptable error cost and a baseline. Collect authorised, representative data; clean missing/invalid values; transform features; split training, validation and test sets; train; evaluate; deploy; monitor and retrain when justified.
  • Fit scalers, imputers and feature selection on training data, then apply the learned transformation to validation/test data. Duplicate entities or future events across splits can create data leakage and unrealistic scores.
  • Overfitting means learning training-specific patterns that generalise poorly; regularisation, representative data and simpler models can help. Underfitting means the model cannot capture useful patterns; improve features or capacity. Hyperparameters control training choices; model parameters are learned.
  • Precision = TP/(TP+FP): of predicted positives, how many are correct? Recall = TP/(TP+FN): of actual positives, how many were found? F1 combines precision and recall. Accuracy can hide failure on a rare class. For regression, error metrics must match the business penalty.
  • Batch inference handles offline collections; real-time endpoints serve interactive calls; asynchronous inference accepts queued work with later results; serverless inference reduces endpoint-management effort for supported patterns. Match latency, payload, traffic and cost.
  • MLOps versions data, code and model artifacts; tracks experiments; automates repeatable pipelines; approves releases; monitors drift and operating health. SageMaker AI supports these lifecycle stages. A technically better score is insufficient if cost per user rises beyond the value of the improvement.

Choose under exam pressure

Requirement Choice and reason
Missing a disease is especially costly Prioritise recall while measuring the resulting false positives.
Many alerts are wrong Investigate precision and the decision threshold.
Millions of records need results tomorrow Evaluate batch inference.

Traps

  • The test set is not a tuning dataset.
  • Data drift does not prove accuracy declined; collect outcome evidence.
  • A model needs operational and business metrics as well as statistical scores.

Active recall

1. TP=80 and FP=20: what is precision?

80/(80+20)=80%.

2. TP=80 and FN=20: what is recall?

80/(80+20)=80%.

3. Why is 99% accuracy weak evidence when only 1% of events are fraud?

Always predicting legitimate already achieves 99%; inspect minority-class performance.

4. Training improves while validation worsens. What is likely?

Overfitting; investigate leakage, regularisation and model complexity.

5. Why retain a model version after replacing it?

For reproducibility, audit and a tested rollback path.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.