Reviewed 10 October 2026. Use the linked official exam guide for your exam version. These are condensed revision notes; the topic pages provide worked distinctions and more recall practice. Google’s 2026 guides use newer Gemini Enterprise Agent Platform names while some APIs and documentation still use Vertex AI.
Memory hook: Data first → baseline → reproducible training → controlled serving → continuous evidence.
Version: follows the published 1 June 2026 scope. It includes conventional ML, foundation models and generative systems. The vendor notes that direct coding skills are not assessed, though interpreting Python/SQL snippets is useful.
1. Low-code and model choice — 1.1–1.2
Use the simplest suitable path: a rule/baseline → pretrained API → BigQuery ML/AutoML → custom training when necessary. BigQuery ML suits SQL-centric data and supported model types; AutoML automates parts of training; Document AI/Vision/Translation provide established capabilities; Model Garden supplies managed/open choices. Evaluate real task quality, cost, latency, privacy, availability, explainability and customization. Gemini handles broad multimodal tasks, Imagen images, Veo video. Supported remote models/tuning differ by model and API.
Classification predicts categories, regression numeric values, forecasting time-dependent outcomes, clustering unlabeled groups. ARIMA-like time-series baselines and interpretable linear/tree models can outperform unnecessary complexity for the business requirement. RAG adds current knowledge; tuning changes behavior; neither removes evaluation.
2. Prepare features and collaborate — 2.1–2.3
Choose SQL/BigQuery, Dataflow, Spark or in-memory Python from scale and complexity. Profile schema, nulls, ranges, labels, sensitive fields and subgroup representation. Split before fitting preprocessing; time-aware/grouped splits prevent future or entity leakage. Features must exist at the real prediction time. A feature store supports consistent definitions and serving, but point-in-time correctness and freshness are still design obligations.
Workbench/Colab notebooks support prototyping; version dependencies/code and avoid embedded credentials. Track dataset snapshot, split, feature transform, code, parameters, hardware, artifacts and metrics in Experiments/Metadata/Registry. A run is reproducible only if its inputs and environment are identifiable.
| Metric | Use and trap |
|---|---|
| Precision = TP/(TP+FP) | Correctness of positive alerts; increase when false alarms are expensive |
| Recall = TP/(TP+FN) | Coverage of real positives; increase when misses are expensive |
| F1 | Harmonic mean of precision/recall; does not encode every business cost |
| MAE / RMSE | Average absolute error / stronger penalty for large errors |
| PR curve / ROC curve | Threshold tradeoffs; PR is often informative for rare positives |
| Calibration | Whether stated probability matches observed outcome frequency |
Generative evaluation adds relevance, groundedness, task completion, safety, latency and cost. Calibrate LLM judges with human-reviewed cases; avoid training/evaluating on the same examples.
3. Scale training — 3.1–3.3
Package deterministic code/environment, validate data and run repeatable jobs. Tune hyperparameters with a bounded trial budget and validation objective; learned parameters are fitted weights, not the search settings. Keep the final test set untouched. Overfitting calls for data/regularization/model changes; input starvation calls for pipeline changes, not indiscriminate accelerators.
CPU fits many preprocessing/conventional tasks, GPU suitable tensor work, TPU supported optimized workloads. Data parallelism replicates the model across batches; model parallelism partitions a model across devices. Communication and memory can dominate. Checkpoint long runs, monitor failed workers/OOM/input bottlenecks, and test restart. Fine-tune only supported models/methods with suitable data and safety checks.
4. Serve and scale — 4.1–4.2
Batch inference meets a completion deadline; online inference meets per-request latency. Version the model and its preprocessing/postprocessing/feature definitions. Managed endpoints reduce operations; Cloud Run/GKE give different execution/control tradeoffs. Choose public/private endpoints, auth, accelerator memory, model-load time, concurrency, quota and feature-store latency deliberately.
Canary limits initial exposure, A/B compares variants, shadow traffic observes new behavior without returning those answers. Scale the full dependency path and retain a compatible rollback model/features. Quantization/compression can lower latency/memory with quality tradeoffs; measure both. Do not assume every deployment supports the same hardware/mode.
5. Automate pipelines and retraining — 5.1–5.2
Ingest → validate → transform → train → evaluate against champion → register → approve/promote → observe. Pipelines/Kubeflow components manage artifacts/dependencies; managed Airflow coordinates broader tasks; Ray serves distributed computation patterns. Reuse preprocessing and record lineage. CI tests code, CD promotes artifacts, CT retrains on time/new data/drift/quality triggers. A retraining trigger is not permission to deploy an inferior model. Cache only when relevant inputs/code match; bound retries and scope pipeline identities.
6. Monitor quality and risks — 6.1–6.2
Data drift changes input distribution; concept drift changes the input/outcome relationship; training-serving skew mismatches training and inference features; attribution drift changes feature influence. Drift is evidence to investigate, not proof of measured quality loss. Labels may arrive late—separate proxy indicators from actual outcomes.
Monitor per-version latency/errors/cost, feature freshness and quality/subgroup metrics. Explainability is not causal proof. Secure data/models/identities, filter sensitive content, test injection/tool abuse and use Model Armor where appropriate. Trace RAG failures to retrieval, prompt, model, tool or policy before deciding to retrain. Roll back or restrict service when the evidence warrants it.
Traps to catch
- Accuracy hides rare-class failure. Random splits can leak time. The test set is not a tuning set.
- Autoscaling inference does not scale every feature dependency. More GPUs can increase communication overhead.
- A completed training job and a passed pipeline do not prove a better production model.
Last-pass self-check
1. Offline quality is excellent but online collapses immediately. First checks?
Feature availability, leakage, preprocessing mismatch, schema/version differences and training-serving skew.
2. Which metric matters when missed fraud is costly?
Recall, considered alongside precision, business cost and an appropriate operating threshold.
3. A model exceeds one accelerator’s memory. What distinction matters?
Model parallelism/memory-efficient techniques versus merely replicating the same full model for data parallelism.
4. Drift alert with no labels: can you claim accuracy fell?
No. Investigate distribution/operational changes and collect outcome evidence; drift alone is not measured accuracy.
5. Does CT mean every retrained model goes live?
No. Validate, compare, apply promotion gates and preserve rollback before deployment.
Sources
- Official exam guide
- Published objective groups (PDF)
- BigQuery ML
- ML pipeline architecture
- Model monitoring
- ML classification metrics
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Choose low-code, APIs or custom ML
Memory hook: Simplest suitable model, measured on the task.
Must remember
- Use BigQuery ML when SQL-centric teams can train and infer near warehouse data; AutoML when supported managed model search meets the task; custom training when architecture or training control demands it.
- Classification predicts categories, regression numeric outcomes, forecasting time-dependent values and clustering unlabeled structure. An interpretable baseline helps expose whether extra complexity is worthwhile.
- Pretrained APIs such as Document AI, Vision and Translation solve established capabilities. Model Garden supplies model options; compare Gemini, image/video models and open models by modality, quality, licensing and deployment needs.
- For generative AI, begin with prompting and grounded retrieval. Fine-tuning can change behavior or domain performance, but it is not the simplest answer to frequently changing facts.
- Evaluate cost per useful result, latency, throughput, privacy, context size and regional availability. A managed model API shifts operations but retains quotas, identity and evaluation responsibilities.
- BigQuery supports ML prediction and supported remote model integration, including supported tuning/inference workflows. Verify the model-specific API rather than assuming every family accepts the same operations.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A SQL analyst needs a tabular baseline | BigQuery ML before building custom distributed infrastructure. |
| An existing document extraction API meets requirements | Use and evaluate that API before training from scratch. |
Traps
- A foundation model benchmark does not establish quality on your own task.
- The most complex model is not automatically the most interpretable or cost-effective.
02 · Data, features and experiments
Memory hook: Same features, known lineage, unseen tests.
Must remember
- Split data before leakage-prone transformations; learn preprocessing parameters only from the training split. Use time-aware splits for forecasting and grouped splits when related records would leak across sets.
- Use BigQuery SQL, Dataflow, Spark or in-memory Python according to scale and transformation needs. Validate schema, nulls, ranges, label quality and distribution before expensive training.
- A feature store helps serve reusable governed features; preserve event time and point-in-time correctness. Shared feature definitions reduce training-serving skew but do not automatically eliminate it.
- Workbench/Colab Enterprise notebooks support exploration. Restrict access, avoid embedded credentials and version code/dependencies; notebooks are not a substitute for reproducible production jobs.
- Track datasets, code, parameters, metrics, artifacts and model versions in experiments and metadata. Compare runs on the same evaluation data and document meaningful changes.
- Use precision/recall, ROC/PR behavior and calibration for appropriate classifiers; MAE/RMSE for regression; task-specific and groundedness/safety evaluation for generation. LLM judges need calibration and human review.
Review details
Metric equations worth remembering: precision = TP/(TP+FP), recall = TP/(TP+FN), and F1 = 2 × precision × recall/(precision+recall). Precision weighs positive-alert correctness; recall weighs missed positives. PR curves often illuminate rare-class behavior better than accuracy. Calibration asks whether a predicted probability matches observed frequency; it is different from ranking performance.
MAE treats errors linearly; RMSE emphasizes larger errors. Choose the metric from business costs and compare at a useful threshold. Feature attribution explains model behavior, not necessarily causation. Group/time-aware splits, fitting transforms on training data only and keeping the final test set untouched prevent misleading scores.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Fraud labels are rare | Inspect precision-recall and business error costs rather than accuracy alone. |
| Offline accuracy is high but online quality collapses | Check feature availability, time leakage and training-serving skew. |
Traps
- Random splits can leak future information in time-series data.
- An LLM judge is an evaluator with limitations, not an infallible ground truth.
03 · Training, tuning and accelerators
Memory hook: Fit the data, distribute the work, checkpoint progress.
Must remember
- Organize training data in suitable storage with repeatable reads and controlled access. Package code, dependencies and configuration into reproducible jobs rather than relying on an interactive notebook state.
- Use supported custom training, AutoML, pipeline components or Kubernetes-based frameworks depending on control needs. Diagnose input starvation, memory pressure, failed workers and incompatible libraries separately.
- Hyperparameters control training choices; learned parameters are fitted from data. Search strategies need a bounded trial budget, objective metric and early stopping where appropriate.
- CPU suits many conventional models and preprocessing; GPUs accelerate suitable tensor workloads; TPUs suit supported accelerator-optimized workloads. Compare utilization, memory and end-to-end time, not just hardware labels.
- Data parallelism distributes batches across model replicas; model parallelism partitions a model that may not fit one device. Distributed training adds communication, synchronization and fault-handling requirements.
- Checkpoint long jobs and retain artifacts/metrics. Fine-tuning foundation models requires compatible data, supported methods and safety/quality evaluation; synthetic data must be checked for errors and bias.
Review details
High training quality with poor validation performance suggests overfitting or data leakage; weak performance on both can suggest underfitting, bad labels or inadequate features. Regularization, simpler models, more representative data and sound validation address different causes. Do not tune repeatedly against the final test set.
Hyperparameter optimization chooses settings such as learning rate or model complexity against a validation objective; learned model parameters are fitted during training. Early stopping and checkpoints bound wasted work. Quantization/compression can reduce serving memory or latency at possible quality cost; benchmark the actual model/runtime before adopting them.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A model exceeds one accelerator’s memory | Consider model partitioning or memory-efficient techniques, not only more independent replicas. |
| Accelerator utilization is low | Investigate input pipeline throughput and CPU/preprocessing bottlenecks. |
Traps
- More GPUs can be slower when communication dominates.
- Tuning against the final test set contaminates the final evaluation.
04 · Serving, rollout and scale
Memory hook: Batch for deadlines; online for requests.
Must remember
- Batch inference processes a collection without per-request interactive latency; online inference serves requests within a response target. Choose the simplest mode that meets the business deadline.
- Register model versions and package compatible prebuilt or custom serving containers. Include preprocessing/postprocessing and feature definitions so online behavior matches evaluation.
- Managed endpoints, Cloud Run and GKE offer different control and operations tradeoffs. Public/private access, authentication, networking and accelerator support must fit the model and consumers.
- Capacity planning includes model load time, memory, concurrency, tokens, request distribution and downstream feature latency. Autoscaling cannot instantly remove cold-start or quota constraints.
- Canary traffic limits exposure to a new version; A/B tests compare variants; shadow traffic can evaluate without returning new-model answers. Keep a tested rollback target and compatible features.
- Feature serving freshness and availability are part of the prediction SLO. Monitor latency percentiles and errors per version, not only average service latency.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Score all accounts before tomorrow morning | Batch prediction may be cheaper and simpler than permanent online capacity. |
| Safely replace a production model | A versioned canary with quality/latency thresholds and rollback. |
Traps
- A model that fits in storage may not fit in serving memory.
- Autoscaling the model does not automatically scale the feature database.
05 · Pipelines and continuous training
Memory hook: Validate before train; evaluate before promote.
Must remember
- An ML pipeline should include ingestion, validation, transformations, training, evaluation, registration and controlled deployment. Version component code, environments and inputs.
- Managed pipelines and Kubeflow-style components express dependencies and artifacts; Airflow coordinates broader workflows; Ray supports distributed computation patterns. Choose by orchestration and execution needs rather than treating them as interchangeable.
- Reuse the same preprocessing specification in training and serving. Data validation catches schema/distribution issues before they become misleading model updates.
- CI tests code/components; CD promotes approved artifacts; continuous training retrains on a defined trigger. A retraining trigger can be time, new data, drift or measured quality degradation.
- Require evaluation gates and lineage before promotion. A freshly trained model may be worse; use a champion/challenger comparison and a rollback plan.
- Use least-privileged pipeline identities, bounded retries, safe caching and artifact retention. Cache reuse is valid only when the relevant inputs and code are equivalent.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| New labelled data arrives weekly | A controlled retraining pipeline with validation and promotion gates. |
| Same preprocessing repeated inconsistently | A versioned reusable component shared by training and serving. |
Traps
- Retraining is not the same as automatically deploying every new model.
- A successful pipeline execution does not prove the model is better.
06 · Monitor quality, safety and change
Memory hook: Data changes; behavior changes; respond with evidence.
Must remember
- Data drift changes input distributions; concept drift changes the relationship between inputs and desired outputs; training-serving skew is a mismatch between training and production features. They require different investigations.
- Model monitoring and continuous evaluation track quality, distributions and feature attribution where supported. Ground-truth labels may arrive late, so use leading indicators without confusing them with proven accuracy.
- Explainability helps understand predictions, but an explanation is not a causal proof. Evaluate performance and fairness across relevant subgroups and document limitations.
- Protect data and model access with identity, encryption, network controls and logging. Guard against prompt injection, sensitive-data disclosure and malicious tool use; Model Armor and safety filters supplement authorization and application validation.
- For generative systems, evaluate retrieval relevance, groundedness, response usefulness, safety, latency and cost. Trace failures to retrieval, prompt, model, tool or policy rather than retraining indiscriminately.
- Set alert thresholds with a response owner and compare model versions. Retrain, roll back, change retrieval or restrict functionality according to the identified cause.
Review details
Feature attribution drift is a change in how features influence predictions; it differs from a change in input distributions. Model monitoring support depends on model/task/version, so do not assume every custom or generative model exposes identical signals. Keep per-version baselines and stratify important user groups.
For delayed labels, monitor schema, missing features, data distributions, latency and user signals as leading indicators, while scheduling true outcome evaluation when labels arrive. Fixing retrieval freshness or a broken feature transform can be the correct response; automatically retraining on every alert may amplify the problem.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Input distribution shifts but labels are delayed | Investigate drift and operational impact while collecting outcome evidence. |
| A RAG answer cites the wrong passage | Inspect retrieval and grounding before assuming the model needs fine-tuning. |
Traps
- Drift is a warning signal, not automatic proof that accuracy has fallen.
- Safety filters cannot authorize a tool action on behalf of a user.