Memory hook: Same features, known lineage, unseen tests.
Must remember
- Split data before leakage-prone transformations; learn preprocessing parameters only from the training split. Use time-aware splits for forecasting and grouped splits when related records would leak across sets.
- Use BigQuery SQL, Dataflow, Spark or in-memory Python according to scale and transformation needs. Validate schema, nulls, ranges, label quality and distribution before expensive training.
- A feature store helps serve reusable governed features; preserve event time and point-in-time correctness. Shared feature definitions reduce training-serving skew but do not automatically eliminate it.
- Workbench/Colab Enterprise notebooks support exploration. Restrict access, avoid embedded credentials and version code/dependencies; notebooks are not a substitute for reproducible production jobs.
- Track datasets, code, parameters, metrics, artifacts and model versions in experiments and metadata. Compare runs on the same evaluation data and document meaningful changes.
- Use precision/recall, ROC/PR behavior and calibration for appropriate classifiers; MAE/RMSE for regression; task-specific and groundedness/safety evaluation for generation. LLM judges need calibration and human review.
Review details
Metric equations worth remembering: precision = TP/(TP+FP), recall = TP/(TP+FN), and F1 = 2 × precision × recall/(precision+recall). Precision weighs positive-alert correctness; recall weighs missed positives. PR curves often illuminate rare-class behavior better than accuracy. Calibration asks whether a predicted probability matches observed frequency; it is different from ranking performance.
MAE treats errors linearly; RMSE emphasizes larger errors. Choose the metric from business costs and compare at a useful threshold. Feature attribution explains model behavior, not necessarily causation. Group/time-aware splits, fitting transforms on training data only and keeping the final test set untouched prevent misleading scores.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Fraud labels are rare | Inspect precision-recall and business error costs rather than accuracy alone. |
| Offline accuracy is high but online quality collapses | Check feature availability, time leakage and training-serving skew. |
Traps
- Random splits can leak future information in time-series data.
- An LLM judge is an evaluator with limitations, not an infallible ground truth.
Active recall
1. What is point-in-time correctness?
Features reflect only information available when the historical prediction would have occurred.
2. Why track dataset versions?
To reproduce a run and explain changes in model behavior.
3. What does a feature store solve?
Managed reusable feature organization/serving, subject to correct freshness and access design.
4. Why pin dependencies?
A different library version can alter preprocessing or model behavior.
5. How compare experiments fairly?
Hold evaluation criteria and data constant while recording changed inputs and parameters.