Reviewed 10 October 2026. Use the linked official exam guide for your exam version. These are condensed revision notes; the topic pages provide worked distinctions and more recall practice. Google’s 2026 guides use newer Gemini Enterprise Agent Platform names while some APIs and documentation still use Vertex AI.
Memory hook: Constraints decide; architecture demonstrates; operations sustain.
Version: current standard scope includes AI/agents. The official case-study set names Altostrat Media, Cymbal Retail, EHR Healthcare and KnightMotives Automotive. Read the linked vendor cases before your sitting; older case-study lists can be stale.
1. Design from requirements — 1.1–1.5
Translate business aims into measurable functional and nonfunctional requirements. Separate mandatory constraints from preferences: residency, latency percentile, availability, RPO/RTO, cost, delivery deadline, skills and license compatibility. Compare build, buy, modify or retire and document the rejected alternative. Six Well-Architected lenses are operations, security, reliability, performance, cost and sustainability.
| Architectural signal | Decision to assess |
|---|---|
| Small team, stateless HTTP/event workload | Managed Cloud Run with bounded scaling and external state |
| Kubernetes-specific control/ecosystem | GKE; Standard/Autopilot by node-control requirements |
| OS or special hardware dependence | Compute Engine; avoid forced containerization |
| Shared POSIX/NFS files / block device / object content | Filestore / supported persistent disk or Hyperdisk / Cloud Storage |
| Familiar OLTP / distributed relational / analytical scans | Cloud SQL or AlloyDB / Spanner / BigQuery |
| Burst absorption and independent consumers | Pub/Sub and idempotent asynchronous processing |
| Private high-throughput hybrid link / encrypted tunnels | Interconnect with explicit encryption / HA VPN |
Multi-zone protects against a zone loss; multi-region only helps when data and dependencies also survive. Queues decouple outages; caches need TTL/invalidation and a source of truth. RPO limits lost data; RTO includes identity, keys, DNS, application reconnection and validation.
Migration sequence: discover dependencies/utilization → disposition and landing zone → test transfer/schema/runtime → baseline+CDC where appropriate → reconcile → cut over under explicit criteria → observe → retire the source after the fallback window. Migration Center supports assessment; a successful copy is not proof of compatible behavior. Calculate transfer time from effective throughput and continuing change rate.
2. Provision the design — 2.1–2.5
VPC global, subnet regional, VM zonal. Shared VPC centralizes networking, peering connects networks without automatic transit, PSC exposes supported services, Private Google Access reaches APIs, NAT handles eligible outbound internet. Select load balancer by protocol, reach, scope and proxy behavior; validate both health-check and data paths. Interconnect does not inherently satisfy encryption requirements.
Use templates/IaC, staged updates, patching and clear controller ownership. Spot fits checkpointed interruption-tolerant work; commitments suit measured steady demand. VMware Engine can fit VMware migration constraints; it is not the default lowest-cost option. Plan storage growth, retention, location and restore—not only initial size.
For AI, compare pretrained APIs, foundation models, AutoML and custom training. Pipelines version preparation → training → evaluation → deployment. Model Garden supplies choices; AI Hypercomputer accelerates suitable jobs. RAG supplies current knowledge, tuning adapts behavior, agent tools perform actions. Search, vision, speech, documents and Gemini Enterprise/NotebookLM workflows solve different user needs.
3. Security and compliance — 3.1–3.2
Apply group/workload identities, short-lived federation, least privilege and separation of duties. Organization policy constrains configurations; IAM grants; applicable deny/PAB enforcement constrains access. IAP and context-aware access protect supported entry points; service perimeters reduce managed-service exfiltration. CMEK adds key control and availability obligations; Secret Manager stores secrets.
Map the actual regulated scope: source data, replicas, logs, backups, exports, operators and AI retrieval. Provider assurance is evidence, not automatic workload compliance. Protect artifact provenance, deployment policy and agent tool authorization; filters alone do not enforce permissions.
4. Optimize technical and business processes — 4.1–4.2
Define SDLC, service catalog, ownership, change approval, test strategy and recovery exercises. Assess team skills and support readiness; a technically sophisticated platform can raise TCO if nobody can operate it. Measure delivery speed, customer outcomes, unit cost and risk. Architecture improvement must follow new evidence and business needs, not feature chasing.
5. Guide implementation — 5.1–5.2
Build once and promote a reviewed artifact through environments. Choose unit, integration, load and security tests for the claim being proved. Use client libraries/CLI with explicit caller/project, enabled APIs, pagination, bounded retries and quota awareness. Emulators test supported behavior but not all production IAM/network/quota controls. Apigee governs API lifecycle; backend authorization still matters.
6. Operate and validate — 6.1–6.6
SLI measures; SLO targets; SLA commits; error budget bounds tolerated unreliability. Alert on actionable user symptoms and budget burn, correlate logs/traces/profiles, and provide runbooks. Canary limits initial traffic, blue/green separates environments, rolling replaces incrementally. Backward-compatible schema is necessary for reliable code rollback. Incidents require mitigation, communication, preserved evidence and blameless improvement. Chaos, penetration and load tests need distinct hypotheses and authorized boundaries.
Traps to catch
- Cheapest VM is not lowest TCO; managed does not mean unmanaged responsibility.
- Regional state defeats a supposedly global design. Replication can copy corruption.
- For case studies, use stated requirements → evidence → choice → tradeoff; do not invent missing constraints.
Last-pass self-check
1. Two answers meet scale. How choose?
Compare mandatory latency, data location, recovery, security, team-operating effort and total cost—not product prestige.
2. What can invalidate a multi-region frontend?
A single-region database, identity/key dependency or untested failover process.
3. What must a case-study decision cite?
The relevant stated business/technical requirement and the evidence showing why the selected design meets it.
4. When does changing DNS fail as a database rollback?
After writes diverge; returning traffic without reconciliation or reverse replication can lose or conflict with new data.
5. Which test proves a backup meets recovery goals?
Restore the data and dependent application under realistic permissions/keys, then measure correctness, RPO and RTO.
Sources
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Projects, Organizations, Billing and CLI
Memory hook: Project contains; billing pays; IAM permits.
Must remember
The resource hierarchy places projects under folders and an organization where available. Projects contain resources, enabled APIs and quotas; billing accounts fund linked projects but are not simply their parent in the IAM hierarchy. Project names, unique IDs and numeric project numbers serve different purposes.
IAM allow policies can be inherited from ancestors; organization policies constrain allowed configurations rather than granting permissions. Cloud Identity/Google Workspace manages organizational identities and groups. Prefer group-based grants to many individual bindings, then review inherited permissions and applicable deny constraints.
Enable required service APIs in the intended project. A quota limits a metric such as resource count or API rate; requesting an increase does not guarantee physical capacity in a specific zone. Select regions for latency, availability, data-location requirements and product support, considering regional versus zonal resource scope.
Budgets alert; they are not automatically a hard spending cap. Export billing data for analysis, use labels and project organization for allocation, and review idle resources, retained disks, public addresses and network transfer. Linking billing and granting access to billing reports require appropriate billing permissions, distinct from workload administration.
gcloud config list and gcloud auth list inspect the active CLI configuration/identity. Named configurations help switch contexts; explicit --project, region and zone flags reduce ambiguity. Cloud Shell supplies a managed command environment but actions still use identity and authorization. Application Default Credentials used by libraries can differ from gcloud's active login: diagnose the actual caller.
Recall drill: explain why a user can administer a VM yet cannot view billing, or can view a project but cannot invoke an API that has not been enabled.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Central identity administration | Cloud Identity/Workspace groups with appropriate IAM grants. |
| Warn at a spend threshold | A billing budget and notifications, plus a separate control process if needed. |
| Prevent prohibited resource configurations | Organization policy constraints. |
Traps
- A budget is not a guaranteed automatic shutdown.
- Changing the gcloud project does not rewrite every script’s explicit project flag.
02 · Compute Engine and Managed Instance Groups
Memory hook: Template repeats; group repairs; disk preserves.
Must remember
Compute Engine provides VMs when OS-level control, custom software or specific hardware is required. Choose machine family, CPU/memory, region/zone, disks, service account and network deliberately. Custom machine types can fit unusual CPU-to-memory ratios. Spot VMs trade lower cost for interruption risk and suit fault-tolerant work.
Boot and data disks have independent lifecycle choices. Persistent disks/Hyperdisk variants have scope, performance and attachment limits; local SSD is ephemeral and should not be the only durable copy. A snapshot supports disk recovery, an image supports reusable boot provisioning, and a machine image captures broader instance configuration/data for supported use cases. Check consistency and retention, not only whether a snapshot job completed.
An instance template defines repeatable VM configuration. A managed instance group (MIG) maintains instances from a template, supports autoscaling, rolling updates and autohealing. Regional MIGs distribute across zones. Load-balancer health checks control traffic; autohealing checks can recreate instances, so thresholds and initial delay should avoid destructive false positives.
Use OS Login for IAM-managed Linux login where supported. Identity-Aware Proxy can provide controlled tunnel access to private instances with the right IAM and firewall configuration. VM Manager supports inventory, patch and OS-configuration operations. Avoid scattering reusable SSH private keys or broad service-account privileges across instances.
Verify inventory, instance status, attached disks, serial/boot logs, service-account permissions and application listeners. Stopping a VM does not necessarily stop all attached-storage or reserved-address charges. Deleting a VM does not guarantee every separately retained disk, snapshot or image disappears.
Review details
Persistent Disk is available in zonal and regional forms; a regional disk replicates across two zones within its region. Replication is not a backup against application deletion. Filestore is managed NFS file storage, Cloud Storage is object storage, and a disk is block storage: choose by application access semantics, not only capacity. Hyperdisk performance/attachment capabilities depend on its type and supported VM.
An unmanaged instance group gathers independently managed VMs for supported use cases; it does not provide a MIG's template-driven desired state and autohealing. A regional MIG still needs a resilient external database/session design. When changing templates, select rolling-update disruption/surge and validate readiness before removing old capacity.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Stateless fleet across zones | Regional MIG, template, health checks and load balancing. |
| Interruptible batch processing | Spot VMs with checkpointing/retry design. |
| Reusable standardized boot environment | A controlled image plus instance template. |
Traps
- Autoscaling, autohealing and load balancing perform different jobs.
- A retained data disk can continue billing after its VM is deleted.
03 · GKE and Container Operations
Memory hook: Google runs the control plane; choose who manages the nodes.
Must remember
GKE Standard gives greater node-pool control; Autopilot manages more infrastructure and applies its supported workload/configuration model. Regional control planes improve control-plane availability; workload replicas and node placement still determine application resilience. Private networking choices govern node addresses and API access separately.
Authenticate to the intended cluster and verify the kubectl context/namespace before changes. Deployments manage stateless replicas and rollouts; StatefulSets manage stable identities and storage patterns; Services discover/expose workloads. Inspect Pods, events, logs, Services and EndpointSlices before changing cluster capacity blindly.
Store images in Artifact Registry and grant the actual pulling identity appropriate access. Image-pull failures can result from a wrong image path, permissions, network restrictions or missing artifacts. Kubernetes ServiceAccounts and Google IAM service accounts are different identities; Workload Identity Federation for GKE connects supported workload identity to Google API access without embedding static keys.
HPA changes Pod replicas; VPA recommends/adjusts resource requests under its configured mode; cluster/node autoscaling changes node capacity. Requests affect scheduling, limits constrain use, and disruption budgets influence voluntary maintenance. Do not expect a Pod autoscaler to manufacture node capacity instantly.
Node pools group node configuration. Plan upgrades, surge/disruption settings, maintenance windows and compatibility. A node count can be healthy while a workload is Pending because of affinity, taints, quota or a PVC. Persistent-volume topology and access modes must match placement.
GKE Enterprise features can help manage fleets, policy and multi-cluster environments. Choose them for actual governance/operational needs rather than assuming every small application requires a fleet platform.
Review details
Probe distinction: readiness removes an unhealthy Pod from eligible service endpoints; liveness restarts a failed container; startup permits slow initialization before normal probes take over. Do not use an aggressive liveness probe to respond to every downstream database timeout.
Read-only diagnosis sequence: kubectl config current-context → kubectl get pods,svc → kubectl describe pod POD → kubectl logs POD. Inspect events for image authorization, scheduling, volume attachment and probe failures. Image pulling often uses a different identity from the application making a Google API call; granting the latter access does not necessarily fix the former.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Minimal node administration | Autopilot when the workload fits its model. |
| Application needs Google API access | Workload Identity Federation with narrowly scoped IAM. |
| Pods Pending after replica increase | Check requests, placement, storage and node capacity. |
Traps
- A regional control plane does not automatically replicate your database.
- Kubernetes RBAC and Google IAM govern different parts of the access path.
04 · Cloud Run, Functions and Event Delivery
Memory hook: Revision receives; identity authorizes; retries repeat.
Must remember
Cloud Run runs containerized services without managing a VM fleet. Services receive requests; Jobs run finite tasks. A revision is an immutable deployment configuration. Split traffic between revisions for controlled release and rollback, and distinguish deploying a revision from sending it production traffic.
Scale and concurrency settings affect latency, cost and downstream connection pressure. Minimum instances reduce cold-start exposure while keeping capacity allocated; maximum instances can protect a backend but do not replace admission control or guarantee unlimited availability. Keep request handlers stateless and externalize durable state.
Cloud Run functions is the current function-oriented experience associated with the older Cloud Functions name. Select generation/runtime/event support deliberately. Eventarc routes supported events to targets; Pub/Sub decouples message producers and subscribers. Cloud Storage object events can trigger processing. Match region, trigger identity, target invocation permissions and event schema.
Design handlers for retries and duplicate delivery. Persist an idempotency key/result or make the operation naturally repeatable. A successful HTTP response acknowledges work; returning success before durable completion can lose business processing. Timeouts and retry policies must match the work and poison-event handling.
Invocation identity and runtime identity are distinct: one calls the service, the other determines what code may access. Restrict ingress and use authenticated invocation where appropriate. VPC egress configuration permits access to private dependencies; it does not automatically make all inbound requests private.
For a rollout issue, inspect the revision receiving traffic, request logs, container startup/listening port, service account, secrets/configuration, concurrency and backend limits. A healthy previous revision offers a rollback option only if data/schema compatibility remains intact.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Request-driven stateless container | Cloud Run service. |
| Finite batch container | Cloud Run Job. |
| Process object-created events | Eventarc/function or Cloud Run handler with idempotent processing. |
Traps
- Deploying a new revision and shifting traffic are separate operations.
- Event-driven does not mean duplicate-free business execution.
05 · Cloud Storage, Transfers and Retention
Memory hook: Object key is not a disk path; retention is not a backup.
Must remember
Cloud Storage stores objects in buckets. Choose location for latency, resilience and data requirements; bucket names and object names are separate identifiers. Object storage is not a normal mutable block filesystem. Version/generation identifiers and preconditions support safe concurrent updates.
| Class | Remember |
|---|---|
| Standard | Frequent access; no minimum storage duration. |
| Nearline | Infrequent access; 30-day minimum storage duration. |
| Coldline | Rare access; 90-day minimum. |
| Archive | Very rare access; 365-day minimum. |
Lower storage price can be offset by retrieval, operations, transfer and early-deletion charges. Archive remains online object storage; do not import another vendor's restore-wait assumptions. Autoclass and lifecycle rules automate supported class/deletion behavior under different models; review compatibility and costs.
Uniform bucket-level access uses IAM rather than per-object ACLs. Public access prevention restricts public grants. Signed URLs provide time-bound access to a specific operation/object under the signing authority; treat the URL as a credential. Encryption is default, while customer-managed keys add key-control and availability responsibilities.
Versioning retains older generations under its model; soft delete supplies a recovery window; retention policies/holds prevent deletion under configured rules. A locked retention policy can be irreversible: understand it conceptually instead of experimenting on a disposable-cost assumption. Lifecycle deletion may be blocked by retention/holds.
Use gcloud storage for object operations, Storage Transfer Service for managed supported transfers and Transfer Appliance for suitable very large offline transfers. Check checksums, permissions, location compatibility and transfer progress. A successful upload is not a tested application restore.
Review details
Cloud Storage provides strong global consistency for object writes, overwrites, deletions and listing. Do not describe object listing as eventually consistent. Access-policy changes take time to propagate, and publicly cached objects can remain stale until their cache lifetime expires: those are different consistency boundaries.
A generation precondition can make a write conditional on the expected object version and prevent lost updates. Single-object operations are atomic, but a batch of several independent object requests is not a multi-object transaction. Pin a generation when a sequence of range reads must use the same version.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Frequently accessed application objects | Standard storage in an appropriate location. |
| Restrict access consistently | Uniform bucket-level access and least-privilege IAM. |
| Bulk recurring transfer | Storage Transfer Service where the source/destination are supported. |
Traps
- Cheap storage classes can incur minimum-duration and retrieval charges.
- A signed URL can be used by whoever possesses it until its conditions expire.
06 · Databases, Analytics and Data Pipelines
Memory hook: Choose by access pattern, not by the largest feature list.
Must remember
| Requirement | Candidate |
|---|---|
| Managed familiar relational engine | Cloud SQL; assess engine, HA, replicas and backups. |
| PostgreSQL-compatible demanding enterprise workload | AlloyDB where its architecture fits. |
| Horizontally scalable relational transactions | Spanner, with deliberate schema/key and location design. |
| Document-oriented application data | Firestore, with document/query/index design. |
| High-throughput wide-column/key access | Bigtable; design row keys to avoid hotspots. |
| Analytical SQL over large datasets | BigQuery; separate from OLTP assumptions. |
| Durable object data | Cloud Storage. |
Availability replicas, read scaling and backup/PITR solve different problems. Verify automatic failover behavior, replication scope/lag and tested restoration for the chosen product. A read replica does not necessarily protect against a destructive write replicated from the primary.
Pub/Sub transports asynchronous messages; Dataflow runs Apache Beam batch/streaming processing; Dataproc runs managed Spark/Hadoop-style workloads. Select by existing code, operational burden, state/window requirements and integration. Inspect job status and logs, not merely the existence of a job resource.
BigQuery organizes projects, datasets, tables and jobs. Partition pruning and clustering can reduce scanned data; selecting needed columns and applying suitable filters controls cost/performance. Load jobs, streaming and external data access differ in latency and economics. bq and SQL query tools operate under the actual caller's IAM permissions and chosen processing location.
Migrate data with supported transfer/replication tools, consistent snapshots or exports, schema conversion where needed and a controlled cutover. Match storage/data-processing locations to avoid unsupported operations, latency or transfer cost. Protect secrets, database network access and service identities; a private IP is not database authorization.
Review details
Cloud SQL regional HA maintains a failover standby with synchronous cross-zone protection; it is different from an asynchronous read replica serving read traffic. Read-after-write through a lagging replica can return old state. Spanner supplies strongly consistent relational transactions; Firestore provides strongly consistent document reads/queries and supported transactions. Bigtable supports strongly consistent reads in a single cluster, while multi-cluster replication/routing can introduce eventual consistency. Choose the documented mode rather than equating “NoSQL” with eventual consistency.
Cloud SQL/AlloyDB auth proxies or connectors simplify supported authenticated TLS connections, but do not create a missing private route. Database users/object privileges are separate from cloud-resource administration. Bound application connection pools across maximum instances.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Global relational consistency and scale | Evaluate Spanner rather than treating all SQL services as interchangeable. |
| Petabyte analytical scans | BigQuery with partition/clustering design. |
| Existing Spark pipeline | Dataproc may reduce migration work. |
| Managed event-time stream processing | Dataflow/Beam with appropriate windows and state. |
Traps
- Bigtable row-key design can create severe hotspots.
- BigQuery is not a direct substitute for a high-frequency transactional database.
07 · VPCs, Subnets and Private Access
Memory hook: VPC is global; subnet is regional; permission is separate.
Must remember
A Google Cloud VPC network is global; its subnets are regional and can serve zones in that region. Custom-mode networks give deliberate address allocation. Plan nonoverlapping primary/secondary ranges for VMs, Pods and services. Supported subnet expansion increases address space; do not assume arbitrary shrinking is available.
Routes decide where packets go; firewall rules/policies decide whether eligible traffic is allowed. Stateful firewall behavior permits matching return traffic, but the route must still exist. Priority, direction, source/destination and target selection matter. Network tags and service-account targets select workloads under their respective rule models; labels used for inventory are not automatically firewall selectors.
Shared VPC centrally manages a network in a host project while service projects deploy authorized resources into it. VPC Network Peering connects networks privately under its route-exchange rules and is not transitive by default. It does not merge IAM policy or all DNS behavior.
Private Google Access lets eligible private-address workloads reach supported Google APIs under correct routing/DNS. Private Service Connect exposes supported services through private endpoints or related service-connectivity patterns. Private services access allocates peering-based connectivity for supported managed services; these similarly named mechanisms are not interchangeable.
Cloud NAT provides configured outbound address translation for eligible private resources without accepting arbitrary unsolicited inbound connections. It needs the relevant routing and regional configuration; it is not an HTTP proxy or a packet-forwarding VM you administer. Reserve static addresses only where stable identity/allowlisting requires them.
Diagnose source identity/address, subnet/range, route, firewall, DNS and destination listener separately. Private networking reduces exposure but does not grant API IAM permissions or database access.
Review details
A new VPC has implied deny ingress and allow egress rules at the lowest effective priority; explicit rules/policies can change the outcome. Lower numeric priority means higher priority within the applicable rule model, but hierarchical/network/VPC policy evaluation must be considered as a whole. The pre-created default network's explicit rules are not the behavior of every custom VPC.
Cloud NAT is regional and uses Cloud Router configuration without routing packets through a Cloud Router VM. NAT port exhaustion can break outbound connections even when the default route exists. Private Google Access also needs appropriate DNS/routes; merely removing a public IP is not a complete private-access design.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Central network across several projects | Shared VPC with controlled service-project attachment. |
| Private workloads need outbound Internet | Appropriate Cloud NAT and routes. |
| Consume supported service at a private endpoint | Private Service Connect. |
Traps
- VPC peering is not automatically transitive.
- A network label is not the same as a firewall network tag.
08 · Load Balancing, Hybrid Connectivity and DNS
Memory hook: Choose protocol, reach and failover scope.
Must remember
Choose load balancing by protocol, external/internal reach, proxy/passthrough behavior and regional/global availability. Application load balancers route HTTP/S; network load balancers serve supported transport-level needs. Backend services, health checks and firewall permissions must agree; a healthy VM does not imply a health-check path is allowed.
Global versus regional resources affect failover and traffic locality. Premium and Standard Network Service Tiers use different network paths and support different product combinations. Check the actual load-balancer type rather than assuming every configuration is global or supports every tier.
Cloud CDN caches eligible content at the edge; cache keys, TTLs and origin protection determine behavior. Cloud Armor applies supported edge/backend security policies. Neither removes the need for secure application authorization or correct cache separation between users.
HA VPN supplies encrypted hybrid tunnels; Cloud Interconnect provides private connectivity with dedicated/partner options. Encryption requirements must be addressed explicitly. Cloud Router manages dynamic BGP route exchange for supported connectivity; it is not itself the data-plane router carrying every packet. Plan redundancy on both provider and on-premises sides.
Cloud DNS public zones publish public records; private zones serve authorized networks. Forwarding and peering zones support hybrid and cross-network resolution patterns. Names resolving correctly does not establish packet reachability. Use split-horizon designs deliberately and avoid forwarding loops.
Reserve static internal/external IPs where a stable endpoint is required. Review health-check sources, firewall rules, backend serving ports, certificates, routing and DNS TTL during migrations. Switching a DNS record does not immediately expire every existing cache or connection.
Review details
For DNS, A/AAAA return IPv4/IPv6 addresses, CNAME aliases a name, MX chooses mail exchangers, TXT carries text data, and NS delegates authority. Public zone creation must be paired with registrar/parent delegation; private zones require authorized networks. DNSSEC authenticates signed answers but does not encrypt queries. TTL controls cache freshness, not a guarantee that every active client connection immediately changes.
For a load-balancer failure, test frontend reachability → selected URL map/backend service → backend health → health-check/data-plane firewall allowance → serving port/application. A proxy load balancer and a passthrough load balancer expose different source-connection behavior; select the precise product before designing source-IP controls.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Private on-premises connectivity | Interconnect with explicit resilience and encryption design. |
| Encrypted hybrid tunnel | HA VPN with redundant BGP/tunnel topology. |
| HTTP host/path routing | Suitable Application Load Balancer configuration. |
Traps
- Cloud Router exchanges routes; it is not the packet-forwarding appliance.
- Private connectivity does not automatically mean encrypted connectivity.
09 · IAM, Service Accounts and Data Security
Memory hook: Grant the action to the actual caller at the narrowest scope.
Must remember
An IAM binding associates a principal with a role on a resource, optionally under conditions. Basic roles are broad; predefined roles are service-oriented; custom roles bundle supported permissions for specific needs. Inherited allow grants can broaden access; deny policies and organization constraints have distinct effects. Removing one narrow grant does not remove an inherited grant elsewhere.
A service account is both an identity used by workloads and a resource whose use can be controlled. Attaching/acting as a service account and creating short-lived impersonated credentials require different permissions. Grant API permissions to the runtime identity, not merely to the human who deployed it.
Prefer attached workload identity, service-account impersonation or Workload Identity Federation over downloaded long-lived keys where supported. External federation exchanges trusted external identity for controlled Google access. Protect audience, attribute mappings/conditions and role bindings; a permissive trust mapping can expose many unintended callers.
Use Secret Manager for secrets and Cloud KMS for encryption-key control. Default encryption does not mean everyone should read the data. Customer-managed keys introduce key IAM, location, rotation, availability and destruction considerations. Separation between key administrators and data users reduces excessive privilege.
VPC Service Controls creates supported service perimeters to reduce data exfiltration; it complements IAM rather than replacing it. Identity-Aware Proxy controls supported application/tunnel access. Audit logs identify activity under their service-specific categories/settings; enable required data-access visibility deliberately and protect the sink destination.
For permission denied, establish the real principal, resource project, required permission, inherited/conditional/deny policies and any perimeter/org-policy restriction. Granting Owner to “make it work” conceals the diagnosis and creates risk.
Review details
Service Account User (roles/iam.serviceAccountUser) includes acting as the service account for supported resource attachment. Service Account Token Creator (roles/iam.serviceAccountTokenCreator) supports generating impersonated credentials and supported signing operations. Granting attachment rights is not the same as giving the human direct access to every resource the account can read, but launching code as that account can be a privilege-escalation path.
Application Default Credentials searches supported credential locations; it is not an IAM role. Workforce federation serves external people, workload federation software. To diagnose a denied call, distinguish authentication failure, missing permission, inherited deny/boundary, organization policy and VPC Service Controls. Principal access boundaries constrain eligible resources for supported access; they do not grant permissions.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Application accesses a bucket | Grant the workload identity the necessary bucket role. |
| CI outside Google Cloud | Federated short-lived credentials with narrowly defined trust. |
| Reduce exfiltration through supported managed APIs | VPC Service Controls plus IAM and data controls. |
Traps
- Service-account use permission is not identical to the service account’s resource permissions.
- A VPC Service Controls perimeter does not replace IAM.
10 · Observability, IaC and Reliable Operations
Memory hook: Observe the user outcome; automate the reviewed intent.
Must remember
Cloud Monitoring handles metrics, dashboards, uptime checks and alerts; Cloud Logging stores/searches log entries and routes them to destinations. Logs-based metrics turn matching entries into measured signals. Custom metrics describe application-specific behavior. Alerts need meaningful thresholds, duration and notification channels.
Log Router sinks select matching entries for destinations such as log buckets, BigQuery, Cloud Storage or Pub/Sub, subject to destination permissions. Retention and exclusions affect cost and future evidence. Creating a sink does not retroactively route all historical logs. Audit-log categories include administration and data access with different defaults/service behavior; verify required visibility.
The Ops Agent collects supported VM telemetry; Managed Service for Prometheus supports Prometheus-style metrics. Traces connect request spans; profiling identifies code resource use. Check provider service status, quota, recent changes and application dependencies while diagnosing.
Infrastructure as code tools include Terraform, Cloud Foundation Toolkit patterns and Config Connector's Kubernetes-style resource management. Helm packages Kubernetes applications. Version configuration, review plans/manifests, protect state and run identities, and verify the resulting service. A controller and Terraform should not compete to manage the same object without a deliberate ownership model.
Use release strategies that match risk: staged/canary traffic, health validation, rollback and backward-compatible schema changes. Define service-level indicators (SLIs), objectives (SLOs) and alerting based on user-visible failure rather than only CPU. Backups need restored-data verification and dependency-aware recovery.
Read-only recall drill: explain what gcloud config list, a filtered Logging query, a Monitoring chart and a Terraform plan each can and cannot prove. No cloud mutation is needed to rehearse the distinctions.
Review details
Remember the four audit categories: Admin Activity, Data Access, System Event and Policy Denied. Admin Activity and System Event logs are always written; most Data Access logs require explicit enablement, while BigQuery Data Access is a notable default-enabled exception. Verify the service, configured exclusions and reader permissions before promising evidence of a past read.
Logging queries filter fields such as resource.type, severity and time; a narrow time/resource scope makes diagnosis easier. Log Analytics supports SQL-based analysis of eligible log buckets. A logs-based metric measures matching entries; an alert still needs an actionable condition, notification destination and owner.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Find a specific failed request | Logs and trace correlation. |
| Detect an error-rate trend | A meaningful metric and alert/SLO condition. |
| Repeat infrastructure deployment | Reviewed IaC with clear resource ownership. |
Traps
- A sink needs destination access and normally routes new matching entries.
- A successful deployment job is not proof of user-visible availability.
11 · Architecture Requirements and Trade-Offs
Memory hook: Requirement first; product second; evidence always.
Must remember
Separate functional requirements (what the system does) from nonfunctional requirements (latency, availability, security, cost and operability). Record constraints, assumptions, dependencies and measurable acceptance criteria. A technically impressive design can fail if the organization lacks the skills or budget to operate it.
The Google Cloud Well-Architected pillars cover operational excellence, security, reliability, performance, cost and sustainability. They are tradeoff lenses rather than six independent checkboxes. Managed services can reduce toil; they still need appropriate identity, data, observability and recovery design.
Choose failure domains based on the required outcome. Multiple zones address zonal failures; a regional dependency can still defeat a multi-region frontend. RPO constrains data loss; RTO restoration time. Replication, backups and failover are different mechanisms. Exercise corrupt data and compromised credentials as well as hardware outages.
Scale stateless work horizontally and externalize necessary state. Cache with an explicit consistency/invalidation strategy. Use queues to absorb bursts and decouple failures, while designing retries, ordering and idempotency. Avoid making every dependency synchronous when a business process tolerates asynchronous completion.
Cost decisions compare total ownership, including people, licenses, transfer, idle capacity and risk. Spot resources fit interruptible work; commitments fit predictable usage only after sizing; serverless fits its execution model rather than guaranteeing lowest cost in every case. Sustainability can align with efficiency and right-sizing, subject to location and service constraints.
For a case study, build a matrix: requirement → evidence in the case → design choice → rejected alternative → remaining risk. Distinguish stated facts from your assumptions. Prefer the least complicated design that meets all mandatory constraints, then plan measured future improvement.
Review details
The currently published standard-exam case-study list is Altostrat Media, Cymbal Retail, EHR Healthcare and KnightMotives Automotive. Use the cases linked from the current exam guide, not an old course's list. Before practice, extract each case's existing estate, explicit business/technical requirements, growth, constraints and key tradeoffs. Do not memorize a fixed “product answer” independent of the scenario.
Storage choice also distinguishes file access (for example, managed NFS through Filestore) from block devices and object APIs. VMware Engine addresses supported VMware migration requirements; it carries a different operating/cost profile from moving a small application to a managed container service.
Current case-study recall
These are original revision cues from the published fictional cases. The design implications are reasoning prompts, not claims about a fixed exam answer.
| Case | Stated constraint to retain | What to reason through |
|---|---|---|
| Altostrat Media | Existing GKE, Cloud Storage, BigQuery and event functions; some ingestion/archive remains on premises; reliability and storage economics matter alongside media AI. | Keep hybrid ingestion and repeatable container operations credible. Match summarization, metadata extraction, recommendations and harmful-content screening to appropriate evaluated capabilities. Auditability and cost remain acceptance criteria. |
| Cymbal Retail | Mixed databases/Kubernetes and legacy file-based integration; supplier content must enrich a product catalog; associates must approve, reject or edit generated content. | Separate extraction/generation, relevant search and human approval before catalog publication. Secure customer interactions and measure discoverability/conversion; do not automate away the explicit review gate. |
| EHR Healthcare | A colocation lease is expiring; containerized customer applications must retain on-premises insurer integrations; existing identity is Active Directory; customer availability must reach at least 99.9%. | Plan hybrid identity/connectivity, consistent container environments, proactive alerts and migration waves. Preserve the stated legacy boundary rather than proposing an immediate rewrite/move of every integration. |
| KnightMotives Automotive | Vehicle/software fragmentation, weak rural connectivity, legacy mainframe/ERP, no dealer hardware budget, prior breaches and EU privacy obligations complicate a five-year modernization. | Stage modernization and skills development, account for connectivity limitations, secure data/ML and dealer workflows, and validate simulation/testing. An always-connected cloud-only assumption or mandatory dealer hardware purchase conflicts with the case. |
Before answering, distinguish a business objective from a mandated mechanism. A case mentioning a technology does not mean it must be replaced, and a growth goal does not override a regulatory or budget constraint.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Strict regional data residency | Keep storage, processing, logs and support dependencies within permitted scope. |
| Small team running variable demand | Prefer suitable managed services and clear operating controls. |
| Two answers both scale | Compare required latency, failure scope, operational burden and cost. |
Traps
- Multi-region frontends do not fix a single-region state bottleneck automatically.
- “Cheapest resource” and “lowest total cost” are different judgments.
12 · Migration, Modernization and Delivery Planning
Memory hook: Discover dependencies before moving the workload.
Must remember
Assess the estate: applications, data, dependencies, owners, licenses, utilization, risk and business criticality. Migration Center supports discovery/assessment. Decide whether to retain, retire, rehost, replatform, refactor or replace each workload according to value and constraints; not every application deserves a rewrite.
Build a landing zone with identity federation, hierarchy, networking, logging, guardrails and billing controls before moving important workloads. Map address overlap, DNS, latency, authentication and firewall dependencies. Shared VPC and hybrid connectivity can support staged migration; they also create transition complexity that must be documented.
Choose online transfer, replication/CDC or offline transfer by dataset size, bandwidth, change rate and downtime tolerance. Estimate transfer time using effective throughput, not nominal link speed, then account for verification and catch-up. Database migration requires engine/schema compatibility, consistency, validation and a defined point at which writes move.
A cutover plan names owners, sequence, prechecks, success criteria, communications, rollback triggers and the latest safe rollback point. DNS changes can leave cached clients on the old path. Dual writes create consistency risks unless deliberately designed; rolling back infrastructure does not magically merge diverged databases.
Modernization choices include containers, managed databases, asynchronous events and serverless services. Preserve observability and operational ownership through the transition. API management such as Apigee can support controlled exposure, quotas and lifecycle, but backend authorization and data correctness remain essential.
Train teams, validate support processes and perform load, security, integration and recovery tests before declaring migration complete. Decommission old infrastructure only after acceptance, retention and rollback obligations are satisfied.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Minimal initial application change | Consider rehosting with a later modernization plan. |
| Low-downtime database transition | Supported replication/CDC plus validated cutover. |
| Unknown application dependencies | Discovery and dependency mapping before scheduling migration waves. |
Traps
- A copied VM is not proof that licenses, identity and latency requirements are met.
- Rollback after new writes requires a data-consistency plan.
13 · AI Platforms, Agents and Responsible Architecture
Memory hook: Ground the answer; constrain the action; measure both.
Must remember
Google's current Gemini Enterprise Agent Platform evolves Vertex AI capabilities into an integrated model/agent platform. Older course and API terminology may still say Vertex AI. Model Garden offers model choices; managed APIs, customization and custom training address different levels of control and effort. Choose a prebuilt API when its capability fits rather than training unnecessarily.
Model choice depends on quality, modality, latency, context, cost, data handling and deployment constraints. RAG retrieves external knowledge at request time; tuning changes model behavior/weights through supported training. Grounding can improve factual relevance but does not guarantee correctness or enforce authorization by itself.
Agent systems add planning, tools, memory and actions. Keep tool authority narrow, authenticate workload identities, validate inputs/outputs and gate high-impact operations. Treat retrieved documents and tool results as untrusted content. Model Armor and Sensitive Data Protection can contribute filtering and sensitive-data controls; application policy and evaluation remain necessary.
ML pipelines orchestrate repeatable data preparation, training, evaluation and deployment. Track datasets, features, experiments and artifacts to reproduce results. Separate training from evaluation data to avoid leakage; monitor production drift and quality. GPUs/TPUs and AI Hypercomputer infrastructure fit different training/serving requirements; expensive hardware alone does not solve bad data or inefficient inference.
Managed search/conversation, vision, document, image, video and audio APIs reduce implementation work for suitable tasks. Gemini Enterprise and NotebookLM-related capabilities support enterprise knowledge workflows under their actual access/governance model. Check supported data locations, quotas, retention and integration rather than assuming all products share one policy.
Evaluate task success, groundedness, safety, latency and cost on representative and adversarial cases. Human review belongs where incorrect output/action has high consequences. Version prompts, retrieval configuration, models and tools as production changes.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Answers need current private knowledge | Permission-aware RAG with evaluation and source grounding. |
| Need a known vision/document capability | Assess a managed API before custom training. |
| Agent can update business records | Scoped tools, authorization, validation and auditable approval boundaries. |
Traps
- RAG does not automatically prevent unauthorized document disclosure.
- An AI-generated infrastructure proposal still needs human/automated verification.
14 · SRE, Release Safety and Operational Excellence
Memory hook: Measure reliability from the user’s side.
Must remember
An SLI measures service behavior, an SLO sets a target and an SLA defines an external commitment/remedy. An error budget is the tolerated unreliability implied by an SLO over its window. Use it to make release/reliability tradeoffs; do not invent a universal acceptable percentage for every business.
Alert on symptoms that require action, with ownership and runbooks. Burn-rate-style alerts compare how quickly the budget is being consumed over suitable windows. Metrics, logs, traces and profiles answer complementary questions. Avoid paging for every transient utilization spike when users are unaffected.
CI validates changes; delivery/deployment moves approved artifacts through environments. Canary releases limit initial exposure; blue/green provides separate environments; rolling updates replace incrementally. Automated rollback needs trustworthy health signals and compatible data/schema. Feature flags separate activation from deployment but require lifecycle cleanup and secure access.
Use load testing for capacity, penetration/security testing for authorized attack resistance and chaos experiments for controlled failure hypotheses. Define blast radius, stop conditions and recovery before experiments. A test in staging may miss production-scale bottlenecks; justify what conclusions it supports.
Incident response should restore service, communicate clearly and preserve enough evidence for a blameless root-cause review. Fix system conditions, not only the person who made the last change. Reduce toil with bounded automation and improve runbooks through actual exercises.
Review capacity, quotas, dependencies, support plans, cost allocation and sustainability continuously. Gemini Cloud Assist and other tools can help investigate or propose changes, but verify recommendations against evidence and the intended environment. Operational excellence is sustained ownership, not a one-time dashboard installation.
Review details
For a request-based SLO, budget is allowed bad-event fraction × valid requests. A 99.9% target across 1,000,000 requests permits 1,000 bad requests. For a time-based 99.9% target over 30 days, the equivalent allowed bad time is 43.2 minutes. Do not mix request and time denominators.
Burn rate is the observed bad-event fraction divided by the budget's allowed bad-event fraction. With a 99.9% objective, a 1% error fraction burns at 10×. Combine suitable short and long windows to distinguish urgent sustained budget consumption from transient noise; the exact alert policy follows the service and response needs.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Rapid releases consume reliability tolerance | Use an explicit error-budget policy and stabilize the service. |
| Risky new version | Canary/staged deployment with meaningful user-health signals. |
| Repeated manual incidents | Automate understood recovery and remove the root cause. |
Traps
- An SLA and an internal SLO can have different targets and consequences.
- Rollback cannot automatically undo an incompatible data migration.