Reviewed 10 October 2026 · Machine Learning Engineer Associate
Memory hook: Prepare without leakage; train with evidence; deploy with rollback; monitor with labels.
Data preparation 28%; model development 24%; deployment/orchestration 24%; operations/security 24%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — Collect and store data
- Choose batch, stream or CDC ingestion from volume, freshness and replay requirements. S3 stores datasets; Glue catalogs/transforms; Kinesis/MSK carry streams; SageMaker AI supports managed ML workflows. Preserve ownership, lineage, versioned source IDs and authorized data access for both traditional ML and foundation models.
1.2 — Perform data transformation, feature engineering, and pre-processing
- Clean missing/invalid values, encode categories, scale where required and transform text/images consistently. Fit transforms on training data only. Feature Store online access serves low-latency features; offline history supports training. Point-in-time joins prevent future facts from leaking into past examples. FM preparation includes deduplication, chunking, embeddings and metadata.
1.3 — Validate data quality and manage bias
- Validate schema, ranges, completeness, duplicates, label quality and train/production representativeness. Split by time, entity or group where random splitting leaks information. Compare subgroup metrics and proxy bias; use Clarify/data-quality tooling as evidence, not a guarantee of fairness.
2.1 — Choose appropriate modeling approaches for ML and AI solutions
- Select classification, regression, clustering, forecasting or an appropriate foundation model based on labels, modality, interpretability, latency and cost. Use pretrained/managed options when they meet requirements. Bedrock simplifies supported FM consumption; SageMaker AI provides custom model lifecycle and hosting control.
2.2 — Train, fine-tune, and customize models for ML and AI solutions
- Version data, code, containers, hyperparameters and seeds. Use validation for tuning and preserve a held-out test set. Match CPU/GPU, input mode and distributed strategy to the bottleneck; checkpoint interruptible training. Compare prompt changes, retrieval, fine-tuning, parameter-efficient adaptation and distillation by what must change.
2.3 — Analyze and evaluate the performance of ML and AI systems
- Precision measures reliability of positive predictions; recall measures found positives; F1 balances both; MAE/RMSE measure numeric errors. Compare training and validation results for over/underfitting. Evaluate FM groundedness, retrieval relevance, safety and actual tool/task outcomes; lower loss or input drift alone does not establish business quality.
3.1 — Manage deployment infrastructure for ML and AI model types
- Real-time endpoints fit interactive latency; asynchronous inference fits queued heavier requests; batch transform fits offline datasets; eligible serverless inference fits intermittent workloads with its feature/cold-start constraints. Multi-model hosting shares infrastructure but introduces loading and isolation trade-offs. Verify model/container compatibility.
3.2 — Provision and configure resources for ML and AI workloads based on existing architecture and requirements
- Provision least-privilege roles, encryption, VPC connectivity, image/model access, quotas and suitable endpoint capacity. Include preprocessing/postprocessing in the deployable contract. Match autoscaling metrics to saturation and startup time; a model artifact alone is not a functioning prediction service.
3.3 — Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads
- SageMaker Pipelines coordinates ML processing/training/evaluation steps; model registry versions and approvals support promotion. CI/CD should publish immutable artifacts, test contracts and gate model quality. Shadow tests compare without serving candidate answers; canary/blue-green controls live exposure with rollback.
4.1 — Monitor ML and AI model inference and performance
- Monitor inference errors, latency percentiles, throughput, utilization and backlog alongside data quality, model quality, bias and feature-attribution drift. Capture data with privacy controls and compare to a baseline. Delayed labels need an outcome-collection process; unexplained input shift is a signal to investigate.
4.2 — Optimize and manage ML and AI infrastructure costs and performance
- Measure total successful-prediction cost. Rightsize, autoscale, batch where suitable, use smaller/compiled/quantized models only after quality testing, and checkpoint eligible Spot training. Identify data-loader, CPU, GPU, network and storage bottlenecks before adding expensive accelerators.
4.3 — Secure ML and AI workloads and model endpoints
- Separate training, inference and pipeline roles; protect model artifacts and datasets with IAM/KMS, TLS and appropriate private connectivity. Container/network isolation, scoped secrets and redacted logs reduce exposure. GenAI workloads also need prompt/tool authorization, tenant boundaries and bounded actions. Govern retraining and rollback rather than blindly retraining on poisoned inputs.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Future data accidentally appears in training | Repair point-in-time features and split discipline. |
| Intermittent inference, cold start acceptable | Evaluate supported serverless options against payload/runtime/features. |
| Need to compare candidate predictions without affecting users | Shadow testing with logged comparable inputs/outcomes. |
| GPU mostly idle during training | Inspect input loading, preprocessing and storage throughput. |
| Production inputs changed but ground truth is late | Report drift and collect outcomes; do not assert an accuracy drop yet. |
Traps
- MLA-C02 includes foundation models and agentic workflows; older MLA-C01-only notes leave gaps.
- Tuning against the final test set contaminates your estimate of generalization.
- A registry entry is not proof of production approval.
- More GPUs and automatic retraining can raise cost or amplify bad data.
Verification cues
- Trace a model version to its training dataset, code/container, evaluation result and approval.
- Inspect CloudWatch endpoint metrics, a Model Monitor baseline/report, pipeline execution and rollback criteria in the console without changing resources.
- Be able to select an inference mode and explain its latency, payload, throughput and operating trade-offs.
Last-pass active recall
1. Why fit normalization only on training data?
To keep validation/test information from leaking into model preparation.
2. What distinguishes data drift from model quality drift?
Input distributions can change without a proven prediction-quality change; quality normally requires outcome labels and suitable metrics.
3. How do you safely use interrupted training?
Checkpoint model/optimizer state and ensure the job can restore within its time constraints.
4. Does successful endpoint deployment prove model quality?
No. Infrastructure health and representative prediction quality require different tests.
5. Why might a small model be preferable?
It may meet quality targets with lower cost and latency; validate the actual task rather than model size.
Sources and version check
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · The ML Lifecycle and MLOps
Memory hook: Split before learning transformations; evaluate before deployment; monitor after deployment.
Must remember
- Start with a measurable business problem, an acceptable error cost and a baseline. Collect authorised, representative data; clean missing/invalid values; transform features; split training, validation and test sets; train; evaluate; deploy; monitor and retrain when justified.
- Fit scalers, imputers and feature selection on training data, then apply the learned transformation to validation/test data. Duplicate entities or future events across splits can create data leakage and unrealistic scores.
- Overfitting means learning training-specific patterns that generalise poorly; regularisation, representative data and simpler models can help. Underfitting means the model cannot capture useful patterns; improve features or capacity. Hyperparameters control training choices; model parameters are learned.
- Precision = TP/(TP+FP): of predicted positives, how many are correct? Recall = TP/(TP+FN): of actual positives, how many were found? F1 combines precision and recall. Accuracy can hide failure on a rare class. For regression, error metrics must match the business penalty.
- Batch inference handles offline collections; real-time endpoints serve interactive calls; asynchronous inference accepts queued work with later results; serverless inference reduces endpoint-management effort for supported patterns. Match latency, payload, traffic and cost.
- MLOps versions data, code and model artifacts; tracks experiments; automates repeatable pipelines; approves releases; monitors drift and operating health. SageMaker AI supports these lifecycle stages. A technically better score is insufficient if cost per user rises beyond the value of the improvement.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Missing a disease is especially costly | Prioritise recall while measuring the resulting false positives. |
| Many alerts are wrong | Investigate precision and the decision threshold. |
| Millions of records need results tomorrow | Evaluate batch inference. |
Traps
- The test set is not a tuning dataset.
- Data drift does not prove accuracy declined; collect outcome evidence.
- A model needs operational and business metrics as well as statistical scores.
02 · Ingestion, Streaming and Transformation
Memory hook: A durable checkpoint and an idempotent sink matter more than a promise of no retries.
Must remember
- Choose batch, micro-batch or streaming using freshness, volume, latency and recovery needs. Full loads copy a dataset; CDC captures changes. Preserve source offsets, timestamps and identifiers so replay, deduplication and audit are possible.
- Kinesis partition keys determine shard placement and per-key ordering; skew can overload a shard. Firehose handles supported delivery/buffering, not arbitrary consumer replay. MSK provides managed Kafka infrastructure; consumer groups and offsets have their own processing semantics.
- Glue jobs and EMR/Spark transform data; Lambda fits bounded event processing; Managed Service for Apache Flink handles stateful streaming. Flink checkpoints preserve recoverable state; event time, processing time, watermarks and late-event policies affect window correctness.
- In Spark, narrow transformations avoid a shuffle; wide operations such as many joins/aggregations redistribute data. Partition skew, tiny files, excessive shuffles and driver collection can dominate runtime. Broadcast a genuinely small dimension when appropriate; do not broadcast a dataset that exhausts executor memory.
- Glue bookmarks track supported processed inputs, not universally exactly-once business output. A failed job can have written partial results. Use transactional tables, staging/commit patterns or idempotent upserts as required.
- Step Functions, Glue workflows or managed Airflow coordinate dependencies and retries according to operational needs. Keep configuration separate from code, package dependencies reproducibly and unit-test transformations plus integration contracts. An SDK paginator is necessary when an API returns continuation tokens.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Correct results for late-arriving event timestamps | Event-time windows with a defined lateness policy. |
| A few keys dominate stream traffic | Reconsider partitioning and downstream ordering requirements. |
| A join shuffles a huge fact table against a tiny dimension | Evaluate a safe broadcast join. |
Traps
- Successful source reads do not prove a committed sink write.
- A bookmark is not a substitute for business-key deduplication.
- Increasing workers may not fix one skewed partition.
03 · Foundation Models and Generative AI
Memory hook: A model predicts plausible output; your application must establish whether it is useful and supported.
Must remember
- A foundation model is pretrained on broad data and can support multiple downstream tasks. Transformers underpin many language models; diffusion models are common for image generation. Multimodal models process or generate more than one modality.
- A token is a model-specific text or data unit, not necessarily a word. A context window limits what a request can include; input and output consume capacity and may have different prices. Longer prompts and answers can increase cost and latency.
- Embeddings represent content as vectors for similarity operations. Chunking divides source content into retrievable pieces. Embeddings do not encrypt text and do not themselves generate a final answer.
- Bedrock offers managed access to supported models and application capabilities. SageMaker AI/JumpStart support a more custom model-development and hosting path. Model choice depends on modality, languages, quality, context/output length, region, licensing, latency, privacy and cost.
- Context engineering selects and organises instructions, retrieved evidence, tool results, conversation state and memory within the available context. Dumping an entire document collection into every request is costly and can obscure relevant facts.
- Generation can hallucinate, vary between calls and be hard to explain. Lower temperature generally reduces sampling variability; it does not guarantee correct or identical answers. Evaluate with real tasks rather than choosing the largest model by default.
- Compare token-based on-demand use, provisioned throughput where supported, caching, batch options and custom-model infrastructure. Business value includes task completion, time saved, user satisfaction and cost per successful interaction, not just tokens per second.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Quickly call several supported foundation models | Bedrock managed model APIs. |
| Need custom training and hosting control | SageMaker AI, with additional operating responsibilities. |
| Relevant facts are missing from the prompt | Improve context/retrieval before buying a larger model. |
Traps
- An embedding is not a lossless copy or a privacy boundary.
- A large context window does not guarantee reliable use of every supplied fact.
- Fluent wording is not evidence of factual correctness.
04 · Prompt Engineering, RAG and Fine-Tuning
Memory hook: Prompt for instructions, retrieve for current evidence, tune for learned behaviour.
Must remember
- A good prompt defines the task, relevant context, output format and constraints. Zero-shot uses no demonstration; one-shot/few-shot include examples. Templates standardise inputs. Version prompts and evaluation sets so improvements can be compared and rolled back; Bedrock Prompt Management supports managed prompt versions.
- Structured reasoning prompts can help decompose a task, but an explanation is not proof of a model's internal process or correctness. Prefer checkable intermediate results and concise justifications. Negative instructions alone are weak enforcement.
- RAG retrieves relevant content, adds it to a prompt and generates a grounded response. A typical path is ingest → clean → chunk → embed → index → retrieve → generate → cite. Bedrock Knowledge Bases manages supported parts of this workflow.
- Vector storage can use supported OpenSearch, Aurora/PostgreSQL or other integrations; the exact service feature matters. Hybrid keyword/vector retrieval, metadata filtering and reranking can improve results. Enforce user access before evidence enters the prompt.
- Fine-tuning changes model weights using task/domain examples; it can improve style and behaviour but is not a live fact lookup. Continued pretraining adapts using additional domain text. Instruction tuning uses instruction/response examples. RLHF uses human preference feedback. Distillation trains a smaller model from a larger model's outputs or behaviour.
- Curate representative, licensed, deduplicated data and hold out evaluation examples. Full pretraining has far higher data/compute requirements than prompt changes or retrieval. Prompt caching can reuse eligible repeated context; it does not repair stale source facts.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Answer questions about policies updated daily | RAG over authorised current documents. |
| Consistent specialised output style across many tasks | Evaluate prompting, then fine-tuning if needed. |
| Reduce a successful model's serving cost | Evaluate a smaller model or distillation against quality targets. |
Traps
- RAG does not alter model weights.
- Fine-tuning does not automatically grant access to new documents.
- A retrieved document can contain hostile instructions; treat it as untrusted data.
05 · Agents, Tools and AWS AI Platforms
Memory hook: The model proposes; tools act; policy decides whether an action is allowed.
Must remember
- An agent uses a model, instructions, state and tools to work toward a goal. A fixed workflow follows predefined steps; an agent can choose among actions dynamically. Use the simpler workflow when its decision structure is sufficient.
- A tool has an interface/schema, permissions and observable results. Validate arguments and outputs. Require human approval for consequential actions where appropriate. Retries need idempotency so a timeout does not create duplicate purchases or updates.
- MCP standardises how supported clients connect to tools and resources; it does not grant trust or make every exposed tool safe. Multi-agent designs may use a supervisor, delegation or peer interaction. More agents introduce coordination, latency and failure modes, not guaranteed quality.
- Short-term state tracks the current task; longer-term memory persists selected facts. Apply data minimisation, isolation, retention and deletion policies. Never mix one customer's retrieved data or memory with another's context.
- Bedrock Agents provides managed agent capabilities; AgentCore provides services for operating agents, including runtime and identity-related capabilities. Strands Agents is an agent-development framework. Choose components based on control and operating requirements, not similar names.
- Current AWS objectives also mention Amazon Quick for business AI experiences and Kiro for AI-assisted development. These are not substitutes for the underlying identity, model evaluation and application security controls. Product branding evolves; use the exam's current terminology.
- Evaluate whole-task completion, valid tool choice, argument accuracy, loop rate, latency and total cost. Add execution limits, safe failure paths and traces so an agent that repeats a tool indefinitely can be diagnosed and stopped.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Known approval steps and deterministic branches | Workflow orchestration. |
| A model must choose among approved business tools | An agent with scoped permissions and validation. |
| Several specialised agents share work | Explicit orchestration, state boundaries and end-to-end evaluation. |
Traps
- Tool discovery is not tool authorisation.
- Memory can preserve incorrect or sensitive information.
- Successful text generation is not the same as successful completion of an external action.
06 · Evaluation and Responsible AI
Memory hook: Measure the answer, the experience and the harm; averages can hide who fails.
Must remember
- Use representative held-out tasks, stable baselines and subgroup analysis. Compare quality, robustness, safety, latency, cost and business outcomes. A benchmark unrelated to the real workflow is weak evidence of suitability.
- BLEU emphasises n-gram precision against references and is associated with translation. ROUGE uses overlap/recall-oriented measures often applied to summaries. BERTScore uses contextual representations for semantic similarity. None alone proves factual correctness or useful business outcomes.
- LLM-as-a-judge can scale evaluation but introduces judge bias, model/version sensitivity and prompt dependence. Calibrate against human judgement and inspect disagreements. Bedrock evaluation capabilities support supported automated and human evaluation approaches.
- Evaluate RAG retrieval separately from answer groundedness and relevance. Evaluate an agent's tool use, action validity and completed task, not just its final wording. Include adversarial inputs and safe refusal/escalation tests.
- Responsible AI includes fairness, robustness, safety, privacy, transparency, explainability, accountability and veracity. Representative data, label review, human audits and subgroup metrics help find harmful differences hidden by aggregate accuracy.
- A transparent system exposes relevant workings and limitations; an explanation helps a person understand a particular result or behaviour. Model cards document intended use, evidence, limitations and risk. An open-source model is not automatically interpretable, safe or appropriately licensed.
- Consider intellectual-property rights, deceptive or biased outputs, environmental impact and user trust. Human-centred design needs clear AI disclosure, feedback and contestability where appropriate. Guardrails reduce specified risks but cannot certify that every response is harmless or true.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Summary wording differs but meaning is similar | Use semantic and human evaluation alongside overlap metrics. |
| High overall accuracy but poor outcomes for one group | Subgroup analysis and fairness investigation. |
| High-consequence decision | Human oversight, documented limits and a suitable error/appeal process. |
Traps
- Fairness has multiple definitions and trade-offs.
- A hallucination score is evidence, not a guarantee.
- Removing all sensitive columns does not necessarily remove proxy bias.
07 · AI Security, Privacy and Governance
Memory hook: Protect the data path and the action path, then keep evidence of both.
Must remember
- Apply least-privilege IAM roles to models, tools, data stores and logs. Separate end-user identity from workload identity. Agent identity and policy features support controls, but the application still needs tenant isolation and scoped tool permissions.
- Use TLS in transit, suitable encryption at rest and controlled KMS key access. PrivateLink provides supported private connectivity; it does not replace identity authorisation. Macie helps discover sensitive S3 data. Secrets should not be included in prompts or source code.
- Prompt injection tries to turn untrusted text into instructions. It may arrive in a user message, retrieved page or tool result. Poisoning corrupts training or indexed data. Jailbreaking seeks to defeat safety behaviour. Validate inputs/outputs, restrict actions and treat retrieved content as data.
- Bedrock Guardrails can apply configured content, topic, sensitive-information and other supported controls. Grounding and output validation can help detect unsupported statements. Do not use model self-reported confidence as the sole authority for high-risk decisions.
- Track provenance, licences and lineage from source data through transformations, model versions and outputs. Define residency, retention and deletion requirements for prompts, logs, embeddings and memory as well as primary datasets.
- CloudTrail supplies supported API audit events; Config evaluates resource configuration; Inspector assesses supported workload vulnerabilities; Artifact supplies AWS compliance evidence; Trusted Advisor highlights supported recommendations. Choose evidence according to the question.
- Governance needs accountable owners, approval gates, review cadence, staff training and documented exceptions. Use a risk framework appropriate to the application and shared-responsibility model. Service compliance does not certify that your own data collection or use is lawful.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A retrieved document tells an agent to reveal secrets | Treat as injection; enforce permissions and tool constraints outside the model. |
| Need evidence of configuration compliance | Config plus the relevant audit process. |
| Need to prove where training material came from | Data lineage, provenance and licensing records. |
Traps
- A private network does not prevent an authorised application from leaking data.
- Encrypting vectors does not resolve the right to retain their source data.
- Logging every prompt without redaction can create a second sensitive-data store.
08 · Feature Engineering, Training and Experimentation
Memory hook: Prevent leakage, track experiments and optimise the metric that represents the real failure cost.
Must remember
- Use Processing/Glue/Spark as appropriate to clean, normalise, encode and transform data. Scale numerical features when the algorithm needs it; categorical encodings, text tokenisation, image augmentation and time-window features solve different problems. Fit transformations only on the training partition.
- A feature store separates reusable feature definitions/values from individual training jobs. Online access serves low-latency features; offline access supports historical training/analysis. Point-in-time correctness prevents future feature values leaking into past examples. Track lineage, schema and training/serving parity.
- Choose algorithms using data type, labels, interpretability, latency and error cost. Handle imbalance with representative sampling, weights or thresholds as appropriate, evaluating against natural production prevalence. Oversampling before splitting can leak duplicate information.
- Version training data, code, container, hyperparameters, seeds and environment. Use experiment tracking to compare trials. Hyperparameter tuning searches choices; validation metrics guide selection while a held-out test set estimates final generalisation.
- Training jobs need suitable CPU/GPU/accelerator, distributed strategy, input mode and storage/network throughput. Checkpoint long jobs; managed Spot training can save cost when interruption recovery and time limits fit. More accelerators can be underused if data loading is the bottleneck.
- For foundation models, compare prompting, retrieval, fine-tuning, continued pretraining and distillation. Parameter-efficient techniques adapt fewer trainable parameters where supported. Evaluate task quality, safety and retrieval/tool performance, not only loss. Document dataset rights and model licence restrictions.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Same features needed for training and serving | Versioned feature pipelines/store with point-in-time correctness. |
| Training repeatedly loses progress on interruptions | Checkpoint and restore model/optimiser state. |
| Validation improves but critical subgroup recall falls | Reject aggregate-only optimisation and investigate subgroup performance. |
Traps
- A random split can leak time or customer identity.
- Low training loss does not prove useful generalisation.
- More GPUs cannot accelerate an input pipeline that cannot feed them.
09 · Model Deployment, Pipelines and Production Monitoring
Memory hook: A model artifact is only one component of a reliable prediction service.
Must remember
- Match real-time, asynchronous, batch or supported serverless inference to traffic, payload, deadline and cold-start tolerance. Multi-model/multi-container deployment options trade consolidation against isolation and operational complexity. Check model/framework compatibility before choosing.
- Package preprocessing, model and postprocessing consistently. Register model versions with evaluation evidence and approval status. SageMaker Pipelines coordinates supported ML steps; CI/CD can promote approved versions through environments using immutable artifacts and scoped roles.
- Use shadow tests to observe a candidate without exposing its answers to users, or canary/blue-green strategies to control live exposure. Define rollback alarms on latency, errors and meaningful model outcomes. A successful endpoint deployment does not prove prediction quality.
- Autoscaling should reflect the actual bottleneck and startup time. Monitor invocation latency/errors, queue depth, utilisation and saturation. Optimisations include batching, model compilation/quantisation where supported, smaller models and suitable inference hardware, each validated for quality loss.
- Model Monitor and related tools help observe data/model quality and drift with appropriate baselines and captured data. Clarify supports bias/explainability capabilities. Delayed ground-truth labels require a separate outcome-collection process; input drift alone is not an accuracy measurement.
- Protect training and inference with separate least-privilege roles, encrypted storage, private connectivity where needed, restricted container/network access and redacted logs. Track prompts/tool calls for AI workloads with privacy-aware retention. Retraining needs governed triggers and approval, not blind reaction to every distribution change.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Evaluate a new model without affecting user responses | Shadow deployment with comparative evaluation. |
| Predictions are slow only after scale-out | Investigate loading/warmup and scaling policy. |
| Inputs shifted but labels are unavailable | Report drift evidence and collect outcomes before asserting accuracy loss. |
Traps
- A registered model is not necessarily approved for production.
- Automatic retraining can amplify bad incoming data.
- Autoscaling capacity is not unlimited or instantaneous.