certslothcertsloth
MLA-C02/Topic 08

AWS / Associate

Feature Engineering, Training and Experimentation

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Prevent leakage, track experiments and optimise the metric that represents the real failure cost.

Must remember

  • Use Processing/Glue/Spark as appropriate to clean, normalise, encode and transform data. Scale numerical features when the algorithm needs it; categorical encodings, text tokenisation, image augmentation and time-window features solve different problems. Fit transformations only on the training partition.
  • A feature store separates reusable feature definitions/values from individual training jobs. Online access serves low-latency features; offline access supports historical training/analysis. Point-in-time correctness prevents future feature values leaking into past examples. Track lineage, schema and training/serving parity.
  • Choose algorithms using data type, labels, interpretability, latency and error cost. Handle imbalance with representative sampling, weights or thresholds as appropriate, evaluating against natural production prevalence. Oversampling before splitting can leak duplicate information.
  • Version training data, code, container, hyperparameters, seeds and environment. Use experiment tracking to compare trials. Hyperparameter tuning searches choices; validation metrics guide selection while a held-out test set estimates final generalisation.
  • Training jobs need suitable CPU/GPU/accelerator, distributed strategy, input mode and storage/network throughput. Checkpoint long jobs; managed Spot training can save cost when interruption recovery and time limits fit. More accelerators can be underused if data loading is the bottleneck.
  • For foundation models, compare prompting, retrieval, fine-tuning, continued pretraining and distillation. Parameter-efficient techniques adapt fewer trainable parameters where supported. Evaluate task quality, safety and retrieval/tool performance, not only loss. Document dataset rights and model licence restrictions.

Choose under exam pressure

Requirement Choice and reason
Same features needed for training and serving Versioned feature pipelines/store with point-in-time correctness.
Training repeatedly loses progress on interruptions Checkpoint and restore model/optimiser state.
Validation improves but critical subgroup recall falls Reject aggregate-only optimisation and investigate subgroup performance.

Traps

  • A random split can leak time or customer identity.
  • Low training loss does not prove useful generalisation.
  • More GPUs cannot accelerate an input pipeline that cannot feed them.

Active recall

1. What is training/serving skew?

The production feature generation or distribution differs from what training used.

2. Why store the exact container version?

Dependencies and runtime changes can alter reproducibility or model behaviour.

3. When is managed Spot training suitable?

When the job can checkpoint, resume and tolerate interruption within its deadline.

4. Why preserve a final held-out test set?

Repeated tuning on it would turn it into another validation set and bias the final estimate.

5. What should be compared before fine-tuning an FM?

A prompt/RAG baseline, task quality, data rights, costs and operational needs.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.