Reviewed 10 October 2026 · DevOps Engineer Professional
Memory hook: Immutable artifacts, constrained automation, observable recovery.
SDLC 22%; IaC 17%; resilience 15%; monitoring 15%; event response 14%; security 17%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — Implement CI/CD pipelines
- CodePipeline coordinates stages; CodeBuild runs buildspec commands; CodeDeploy controls supported deployments. Separate build/test/deploy roles and environments. Cross-account pipelines need role trust, artifact-bucket access and KMS permissions; untrusted pull requests must not inherit production credentials.
1.2 — Integrate automated testing into CI/CD pipelines
- Combine unit, contract, integration, security and end-to-end tests. Promote only reviewed artifacts after acceptance gates; mocked services do not validate real IAM/network behavior. Manual approval is useful only when it shows the exact artifact and evidence being approved.
1.3 — Build and manage artifacts
- Version source, dependencies, build tools and artifacts. ECR image digests identify content while tags can move; CodeArtifact serves packages. Scan/sign where required, record provenance and keep secrets out of layers/logs. Build once and promote the same artifact.
1.4 — Implement deployment strategies for instance, container, and serverless environments
- Compare in-place, rolling, immutable, blue/green, canary and linear delivery. Lambda versions/aliases support eligible weighted rollout; ECS/EC2 need deployment-specific traffic/health hooks and capacity. Alarms and validation can trigger rollback, but destructive database changes need separate compatible sequencing.
2.1 — Define cloud infrastructure and reusable components to provision and manage systems throughout their lifecycle
- CloudFormation/CDK/SAM describe repeatable infrastructure. Parameters vary input; mappings select fixed values; conditions choose branches; outputs expose results. Understand dependencies, change sets, drift, replacement, rollback and retention. Creation signals test readiness rather than merely instance-running state.
2.2 — Deploy automation to create, onboard, and secure AWS accounts in a multi-account or multi-Region environment
- Organizations, Control Tower, StackSets and delegated administration support account baselines at scale. Scope automation roles, SCPs and boundaries; grant only necessary service integration. Account creation success does not prove logging, networking, cost ownership and emergency access are ready.
2.3 — Design and build automated solutions for complex tasks and large-scale environments
- Systems Manager Automation coordinates runbooks; State Manager maintains associations; Run Command executes actions; Patch Manager uses baselines and maintenance windows. EventBridge and Step Functions can coordinate fleet work. Bound concurrency/errors, validate parameters and make replays idempotent.
3.1 — Implement highly available solutions to meet resilience and business requirements
- Use multi-AZ useful capacity, health checks, stateless compute and redundant dependencies. Match RDS/Aurora deployment behavior to failover/read requirements. Eliminate shared single points in NAT, DNS and data access; fault isolation is more than duplicating web servers.
3.2 — Implement solutions that are scalable to meet business requirements
- Scale on meaningful demand: CPU for CPU-bound work, requests or backlog per worker when appropriate. Account for warmup, lifecycle hooks, quotas, downstream capacity and graceful draining. Scheduled/predictive scaling anticipates load; target tracking and step scaling react to measured demand.
3.3 — Implement automated recovery processes to meet RTO and RPO requirements
- Define RPO/RTO and automate tested backups, replacement and regional recovery. Include artifacts, keys, secrets, identity, quotas and data reconciliation. Replica promotion and traffic movement must respect write ownership; plan failback and measure restored business behavior.
4.1 — Configure the collection, aggregation, and storage of logs and metrics
- Collect structured application/system logs, metrics and traces centrally with least privilege and retention. CloudWatch agents add guest metrics/logs; native EC2 metrics do not include every OS measurement. Log subscriptions, metric streams and exports are different delivery paths with distinct permissions.
4.2 — Audit, monitor, and analyze logs and metrics to detect issues
- CloudTrail answers API activity; Config answers supported configuration/compliance; CloudWatch and traces answer health and latency. Use Logs Insights, correlation IDs and baseline/anomaly signals. Protect audit destinations from workload-account deletion; enable necessary data events rather than assuming all are present.
4.3 — Automate monitoring and event management of complex environments
- Use metric/composite alarms, EventBridge event patterns/schedules, buses and target roles deliberately. Add retries, dead-letter capture, deduplication and end-to-end observability. Missing expected events can matter as much as explicit failure events.
5.1 — Manage event sources to process, notify, and take action in response to events
- Identify the actual event source and delivery semantics. SQS visibility/redrive, Lambda asynchronous destinations, stream checkpoints and EventBridge retry/DLQ behavior are not interchangeable. Preserve correlation and operation IDs so replay does not duplicate side effects.
5.2 — Implement configuration changes in response to events
- Remediation needs current-state checks, bounded permissions, safe parameters, rate limits and rollback. A Config fix must not fight the deployment system in an endless loop. Route sensitive or uncertain actions to review; verify the postcondition before declaring success.
5.3 — Troubleshoot system and application failures
- Diagnose earliest meaningful pipeline/stack failure, then caller/action/policy, DNS/routes/security rules, quotas and dependency health. Distinguish throttle, timeout and invalid input. Use traces and latency percentiles; adding retries can deepen overload.
6.1 — Implement techniques for identity and access management at scale
- Use federation and short-lived roles; constrain trust, iam:PassRole, boundaries and SCPs. Separate human, build and deployment identities. Resource policies and KMS key policies must support cross-account operations; centralized billing is not an access grant.
6.2 — Apply automation for security controls and data protection
- Automate encryption, secret retrieval/rotation, image/dependency assessment and configuration baselines. Apply preventive checks before deployment and detective/remediation controls after it. Protect pipeline credentials and supply-chain integrity; a passing scanner is not proof an artifact is harmless.
6.3 — Implement security monitoring and auditing solutions
- Aggregate GuardDuty, Inspector, Security Hub, CloudTrail and Config evidence according to the requirement. Test incident runbooks, isolation, credential containment, evidence preservation and recovery. Review incidents to improve controls and service objectives rather than merely adding an alarm.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Release must stop when error rate rises | Canary/linear rollout with meaningful alarms and rollback. |
| Same baseline in many accounts/Regions | StackSets with appropriate delegated permissions. |
| Unexpected console edits | Drift detection for supported properties; repair deliberately. |
| Fleet repair starts repeatedly | Idempotency, current-state checks and bounded retry/error policy. |
| Cross-account artifact decryption fails | Inspect KMS grants/key policy as well as S3 and role trust. |
Traps
- Infrastructure success does not prove application readiness.
- Code rollback does not automatically reverse data mutations.
- Automation can amplify incidents when scope and error controls are missing.
- A queue backlog is not always solved by more workers; downstream capacity may be saturated.
Verification cues
- Read a buildspec, deployment AppSpec/lifecycle hooks, CloudFormation change set and earliest stack failure event.
- Trace a release from source revision to artifact digest, approval, deployment and rollback alarm.
- For a remediation event, explain source policy, target role, retry/DLQ handling, runbook scope and verified outcome.
Last-pass active recall
1. Why promote instead of rebuilding?
Rebuilding can change dependencies/content; promotion preserves the tested artifact.
2. What separates a boundary from an identity policy?
A boundary constrains applicable permissions; it grants none by itself.
3. Why are database migrations often expand/migrate/contract?
Old and new code need a compatible transition and viable rollback window.
4. What does a successful recovery drill measure?
Useful business service, data loss/correctness and elapsed restoration time with all dependencies.
5. Why can archived-event replay be dangerous?
It repeats delivery and may repeat side effects unless consumers are idempotent and scope is controlled.
Sources and version check
- Official DOP-C02 exam guide
- CodeDeploy deployments
- CloudFormation concepts
- Systems Manager Automation
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Artifacts, SAM and Continuous Delivery
Memory hook: Build once, identify the artifact, test it, then move traffic deliberately.
Must remember
- Keep source, dependencies, build instructions and infrastructure definitions versioned. CodeBuild executes a buildspec; CodePipeline coordinates stages and artifacts; CodeDeploy manages supported deployment strategies. Store immutable, identifiable artifacts in suitable S3/ECR/package repositories.
- A Lambda deployment package must match its runtime and CPU architecture. Layers share supported dependencies; container images use a compatible runtime interface. An image tag can move: an immutable digest identifies what was actually deployed. Never bake runtime secrets into an artifact.
- SAM expresses serverless infrastructure using a CloudFormation transform.
sam buildprepares artifacts; local invocation supports development feedback but cannot reproduce every cloud permission, network or managed-service behaviour. CloudFormation change sets describe proposed changes; they do not prove application correctness. - Test in layers: unit tests for logic, integration tests for actual service contracts, contract/schema tests for producer/consumer compatibility and end-to-end smoke tests for useful service. Mocking a success response does not validate IAM or eventual consistency.
- Versions publish immutable Lambda configurations/code; aliases point to versions and support eligible traffic splitting. Canary exposes a small proportion first; linear shifts traffic in increments; all-at-once moves it together. Use alarms and lifecycle validation hooks to stop/roll back unsafe releases.
- ECS blue/green deployment uses separate task sets/traffic destinations for supported controllers. EC2 deployments need healthy spare capacity and suitable hooks. Database changes require backward-compatible sequencing because rolling code back may not reverse a destructive schema change.
- Separate development/staging/production roles and configuration. Approvals gate risk; automated tests detect behaviour. A successful pipeline means its configured steps passed, not that every failure mode has been tested.
- AppConfig separates configuration and feature-flag release from application binaries. Validators check proposed configuration, deployment strategies control rollout, and configured alarms support rollback. A configuration change can break the application without any code deployment, so test compatibility and choose safe defaults.
- Current emerging-topic check (10 October 2026): the official DVA-C02 guide lists broader AI-assisted code generation/review, test generation, CI/CD assistance, troubleshooting and optimisation as possible unscored pretest topics; Amazon Q Developer assistance also appears within the published development-domain skills. Treat generated code, tests and repair suggestions as proposals requiring review and representative validation. Keep secrets and personal data out of inappropriate model inputs/logs; scope agent tools and preserve explicit deployment authority. The emerging-topic list is not an additional weighted exam domain.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Release gradually and stop on error-rate increase | Canary/linear delivery with alarms and rollback. |
| Verify IAM against an actual database | Integration test in an isolated environment. |
| Ensure production runs the reviewed container | Pin and record the image digest. |
Traps
- Rebuilding separately in production may produce a different artifact.
- A change set previews infrastructure actions, not test outcomes.
- Rollback requires a compatible data and dependency state.
02 · CloudFormation and Systems Manager Operations
Memory hook: Inspect the intended change, the actual state and the identity performing it.
Must remember
- A CloudFormation stack tracks declared resources; dependencies control ordering. Change sets preview updates, drift detection compares supported properties with actual state, and StackSets apply stacks across selected accounts/Regions. These are different operations.
- Know update-in-place versus replacement, rollback states and why a retained resource can survive stack deletion.
DeletionPolicycontrols supported deletion/retention behaviour; update-replacement retention is a separate concern. Imported resources and existing physical names need careful ownership checks. - Template parameters vary inputs; mappings select fixed values; conditions select resources/properties; outputs expose results. Resolve circular dependencies by reconsidering resource references and ordering. A creation signal or wait condition is not automatically satisfied by an EC2 instance entering running state.
- Use service roles with least privilege and understand
iam:PassRole. A user may initiate an operation while a service role performs it. Read stack events from the earliest meaningful failure, not only the final rollback summary. - Systems Manager managed nodes need a working agent, identity permissions and connectivity to required endpoints. Session Manager avoids opening SSH/RDP ports for supported access. Run Command executes commands, State Manager maintains associations, Automation coordinates runbooks, Patch Manager applies patch policies.
- Maintenance windows define when approved work may run; patch baselines define approved patches. Inventory and compliance show state; remediation requires a configured action. Restrict runbook parameters and use approvals, rate controls and failure thresholds where the impact demands them.
- Schedule start/stop and cleanup by explicit ownership tags. Stopped compute can leave charged storage and addresses. Cost Explorer, budgets, tags and anomaly detection provide evidence and alerts, not immediate universal spending caps.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Same declared baseline across many accounts | StackSets with appropriate delegated permissions. |
| Find console edits to managed resources | Drift detection for supported properties. |
| Run an approved multi-step repair | Systems Manager Automation runbook. |
Traps
- Drift detection does not automatically repair drift.
- A successful stack update does not prove application readiness.
- A service role can perform actions the initiating user cannot perform directly; control who may pass it.
03 · ELB & Auto Scaling
Memory hook: A load balancer chooses a destination; an Auto Scaling group maintains capacity; application state must survive either decision.
Must remember
Pick the right traffic layer
- Application Load Balancer (ALB) understands HTTP/HTTPS at layer 7. A listener receives traffic; ordered rules choose forward, redirect, authentication or fixed-response actions. Conditions can include host and path.
- Target groups hold instances, IPs or supported Lambda destinations, with health checks/attributes. Path rules can select different groups.
- Lower numbered listener priorities run first. A fixed response comes from the ALB and can succeed even when the application is unavailable.
- Network Load Balancer (NLB) handles TCP/UDP/TLS at layer 4 and supplies static addresses per enabled AZ, with optional Elastic IPs for supported internet-facing deployments. It fits non-HTTP traffic and IP allowlist requirements.
- NLB is not an HTTP path router. Modern NLBs support security groups in supported configurations; the old claim that NLBs never have security groups is unsafe.
- Gateway Load Balancer (GWLB) steers traffic through network appliances using GENEVE/UDP 6081, rather than routing application URLs.
- An internet-facing load balancer can forward to private backends; public users do not require public backend IPv4.
Connection behavior and health
- Stickiness uses supported cookies/affinity mechanisms to favor a target. It does not replicate session memory or guarantee that target will survive.
- Cross-zone load balancing allows a node to route across enabled AZs. ALB has it enabled at the load-balancer level, with target-group-level controls; NLB/GWLB default differently. Check transfer cost rules for the specific type.
- Deregistration delay drains in-flight requests during removal. Too little can interrupt long work; excessive values can slow scale-in and deployments.
- TLS termination uses listener certificates, often from ACM. SNI lets a compatible client indicate its hostname so a listener chooses the correct certificate. ALB certificates are regional; CloudFront ACM certificates use us-east-1.
- Target health checks detect application reachability on the configured port/path. EC2 status health and application health answer different questions.
- Health routing is not authorization: all-unhealthy/fail-open behavior can send traffic to unhealthy targets.
Scaling and replacement
- Vertical scaling changes machine size and may require interruption; horizontal scaling changes the number of workers. Externalized state makes horizontal replacement safer.
- AWS Auto Scaling plans coordinate resources. EC2 Auto Scaling manages instance groups; Application Auto Scaling handles supported dimensions such as ECS task counts. Plans can migrate to direct policies.
- Launch templates describe AMIs, type, user data, interfaces and related launch settings. Updating a template does not automatically update every already-running instance.
- An ASG maintains desired capacity between minimum and maximum bounds. Enable appropriate ELB health when application failures should cause replacement; EC2-only checks may miss a broken web process.
- Instance refresh rolls out launch changes with health, warmup and capacity constraints.
- Target tracking maintains a chosen metric target; step scaling changes capacity according to alarm severity; scheduled scaling anticipates known times; predictive scaling forecasts recurring demand.
- Choose a demand-related metric: CPU can fit compute-bound servers, ALB request count per target can fit web workers, and queue backlog per worker can fit asynchronous processing.
- Warmup/cooldown reduce unstable decisions during startup; grace periods do not prove readiness.
- Multi-AZ subnets with desired capacity one do not provide two active replicas. Production resilience needs sufficient surviving capacity and dependencies.
Choose under exam pressure
| Requirement | Decision and reason |
|---|---|
| Several web apps under paths or hostnames | ALB listener rules and target groups |
| Static addresses for TCP/UDP clients | NLB |
| Third-party network inspection fleet | GWLB |
| Scale before a known daily opening | Scheduled scaling |
| Maintain utilization near a target | Target tracking |
| Replace instances with a broken web process | ASG using relevant ELB health |
| Roll out a new AMI to existing capacity | Controlled instance refresh |
Traps
- Scaling and healing differ. Replacing an unhealthy instance can preserve the same desired capacity without adding demand capacity.
- A cookie is not a session database. Target loss still destroys target-local state.
- A successful listener test can bypass dependencies. Fixed responses do not prove backend, cache or database health.
04 · Disaster Recovery & Migrations
Memory hook: RPO limits how much data may be lost; RTO limits how long useful service may be unavailable.
Must remember
Objectives before recovery patterns
- Recovery point objective (RPO) measures acceptable data loss as a time interval. Recovery time objective (RTO) measures acceptable recovery duration. Fast restoration cannot recreate transactions that never reached the available recovery data.
- Measure detection, decision, provisioning, data recovery, dependencies, traffic switching and validation. A database restore benchmark alone does not prove users can work within the RTO.
- Backup and restore: retain recovery data, then rebuild as needed. Pilot light: keep the minimal core running and add missing serving capacity. Warm standby: maintain a functional reduced-capacity workload and scale it. Multi-site active: serve from multiple sites while managing routing, consistency and failure behavior.
- Recovery patterns trade standing cost and complexity against recovery work; their names do not guarantee an RPO/RTO. Select using tested behavior and actual requirements.
- Availability across AZs and regional disaster recovery solve different failure scopes. Multi-AZ replication does not automatically protect against a Region-wide outage, accidental deletion or corrupted application data.
Backups need more than a schedule
- AWS Backup plans define rules; resource selections determine what is protected; vaults contain recovery points. A plan with no selections schedules no useful protection for the intended resource.
- Cross-Region or cross-account copies require compatible resource support, IAM and encryption-key arrangements. Verify the copy, its retention and the destination's ability to restore. A copied recovery point that cannot be decrypted is not a working recovery plan.
- Vault Lock provides retention controls; compliance-mode protection can intentionally prevent early deletion. Governance controls and compliance immutability have different bypass properties. This pack does not create retention that obstructs immediate cleanup.
- Keep historical recovery points where required: replication can quickly copy an unwanted deletion or bad write. Backup retention, replication and application validation protect against different failure modes.
- Validate service quotas, instance availability, subnet addresses, dependencies and key access in the standby Region before a disaster. A successful Terraform plan does not reserve all future capacity or raise every quota.
Migrate with a deliberate cutover
- DMS supports data movement with full load and change data capture (CDC) for supported endpoints. SCT helps convert schema/code for heterogeneous migrations; data replication does not automatically convert every stored procedure or engine feature.
- A low-downtime database migration commonly uses initial load, ongoing CDC, validation, a controlled write cutover and a rollback decision. Monitor replication lag and reconcile data; “CDC enabled” is not proof of zero data loss.
- RDS/Aurora paths include compatible dump/restore, snapshots, replicas and DMS. Check engine/version support and downtime tolerance before selecting a path.
- VM Import/Export moves supported VM images. Application Migration Service (MGN), now documented as AWS Transform MGN, replicates servers for rehosting and cutover, with staging resources and costs. Rehosting differs from moving only a database.
- Application Discovery Service is a historical inventory/discovery tool closed to new customers, not a universal replacement for MGN. VMware Cloud on AWS preserves VMware-based operating assumptions but requires current commercial availability and capacity checks. It remains named in the published exam list; that does not authorize a new subscription in this lab.
- For large datasets, estimate effective transfer time from data size and throughput, including validation and changes during transfer. Compare DataSync, suitable uploads and supported alternatives; transfer/storage notes cover the tool boundaries. Snow Family remains a published exam concept despite lifecycle changes; no physical job is created here. Consult service status.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Restore rarely, tolerate substantial recovery work | Backup and restore |
| Keep critical data/core services ready but provision serving tiers later | Pilot light |
| Need a functional secondary environment before failure | Warm standby |
| Minimize database migration downtime | Full load plus CDC, validation and controlled cutover |
| Rehost whole servers with limited rewriting | MGN |
| Convert between database engine families | Schema assessment/conversion plus a compatible data migration |
| Centralize protection across supported resources | AWS Backup plans, selections and restore tests |
| A standby cannot scale during disaster | Inspect quotas, capacity, address space and dependencies |
Traps
- Faster compute during restore improves no data that is absent from the backup.
- A replica is not automatically a historical backup, and a backup is not automatically immediately available serving capacity.
- Retention locks deliberately constrain deletion. Terraform lifecycle flags cannot bypass a genuine compliance retention requirement.
05 · Monitoring & Audit
Memory hook: Metrics show symptoms, logs explain events, traces follow requests, and audit records identify changes.
Must remember
Separate the evidence questions
- CloudWatch: how is the workload behaving? Use metrics, logs, dashboards and alarms for operational evidence. CloudTrail: which identity called which API, against what resource and when? AWS Config: what resource configuration existed, and did an evaluated rule consider it compliant?
- These sources complement each other. A slow API can require a latency alarm, application logs and a trace; identifying an administrator's change requires audit evidence. Check that the needed events, resources and retention were actually configured before promising historical answers.
- AWS X-Ray follows instrumented requests across services and downstream calls. Its trace map helps find latency, errors and bottlenecks. Traces are not a replacement for every application log or API audit event; instrumentation and sampling affect visibility.
Metrics and logs
- A metric is a numeric time series identified by namespace, name and dimensions. Choose meaningful statistics and evaluation periods: average latency can conceal slow tail requests, while a total error count without request volume may mislead.
- Alarms evaluate metric conditions; actions notify or invoke supported responses. Composite alarms combine alarm states to reduce noisy paging. Treat missing data intentionally rather than assuming missing means healthy. Metric streams continuously deliver selected metric updates to downstream consumers.
- Logs Insights queries log events. Metric filters count matching events into metrics. Subscriptions forward matching logs to supported destinations; export writes log data to S3 for a different processing workflow. These are distinct mechanisms.
- The CloudWatch agent collects additional guest-OS and application signals. Standard EC2 metrics do not automatically reveal every filesystem or memory measurement.
- Container Insights and Lambda Insights add workload-specific visibility; Contributor Insights identifies prominent contributors; Application Insights helps correlate application problems. Their scope and collection costs need deliberate configuration.
Events, audit and compliance
- EventBridge rules match events on buses and deliver them to targets. Scheduler invokes targets on time-based schedules. An archive retains selected events for replay; a replay can repeat a business action, so consumers still need idempotency.
- A successfully accepted event can fail to match a rule or fail later delivery. Check source/detail pattern, bus, target configuration, resource permissions and retry/dead-letter behavior separately.
- CloudTrail management events describe control-plane activity; selected data events provide supported resource-level activity. CloudTrail Insights detects unusual supported API activity. A rule reacting to an API call via EventBridge needs the appropriate event path and coverage.
- Config recording tracks selected resource configurations; rules evaluate compliance. Notifications and remediation are separately configured. A remediation role can mutate resources, so evaluation should not be confused with automatic repair.
Additional published-scope tools
- Amazon Managed Service for Prometheus stores and queries compatible operational metrics, especially for container workloads using PromQL. Amazon Managed Grafana visualizes metrics, logs and traces from multiple sources. The dashboard layer is different from the metric storage/query layer.
- AWS Health Dashboard reports AWS service events and account-relevant impacts. Combine it with workload telemetry: a healthy AWS status does not prove your application or configuration is healthy.
- Keep logs and traces useful: redact sensitive data, set retention, scope collection and correlate request identifiers. Broad logging can create both sensitive-data exposure and substantial ingestion charges.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Alert on errors or latency | CloudWatch metric alarm with an action |
| Determine who deleted a database | CloudTrail audit events |
| Review a resource's historical configuration compliance | Config history and rule evaluations |
| Find the slow downstream call in a distributed request | X-Ray tracing |
| Query container metrics using PromQL | Managed Service for Prometheus |
| Visualize several telemetry sources together | Managed Grafana |
| Trigger work for matching application events | EventBridge rule and target |
| Investigate a relevant AWS service disruption | AWS Health Dashboard |
Traps
- An alarm with no action is not an email subscription; rule creation alone does not establish every permission needed for delivery.
- A trace sample or a log metric is not a complete security audit record.
- A Config rule can report a problem without correcting it. Replaying an event can repeat side effects rather than merely replaying a picture of history.
06 · Policy Evaluation and Security at Scale
Memory hook: A grant, a ceiling, a trust relationship and a network path are four separate checks.
Must remember
- Identify the actual principal ARN/session, requested action, resource and condition keys. Identity policies and resource policies can grant access under different rules; boundaries, session policies and SCPs constrain applicable permissions. An explicit deny dominates. Role-session and same-account resource-policy exceptions mean simplistic intersection slogans can mislead.
- Cross-account role access needs a trusting target role and permission for the caller to assume it, plus no applicable deny. External IDs help prevent a third-party confused-deputy problem; they are not passwords. Use source-account/source-ARN conditions where supported for AWS service access.
- IAM Identity Center permission sets support workforce account access. Federation trust, session duration, MFA and emergency access need deliberate design. Permission boundaries delegate role creation while limiting potential rights; control who may change or remove the boundary.
- KMS has its own key-policy/grant and identity-permission evaluation. Cross-account encrypted-data access needs both data-resource access and the relevant key permissions. Rotating a key does not re-encrypt every existing object automatically. Multi-Region related keys do not make all policy and resource configuration global.
- SCPs, resource control policies where supported, organisation conditions and account boundaries serve different purposes. Apply controls through OUs and delegated administration, test exceptions and keep an audited emergency path. A deny aimed at one Region must account for global service behaviour.
- Use infrastructure pipelines, Config rules, security standards and automated remediation to maintain baselines. Classify controls as preventive, detective or responsive. Record evidence and exceptions; compliance frameworks guide controls but do not replace risk analysis.
- For containers and serverless, separate runtime role, execution/image-pull role and infrastructure role. Scan images/dependencies, restrict network paths and secrets, and audit tool/service access. For AI, treat prompts/retrieved content as untrusted and authorise tool actions outside the model.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Delegate role creation without unlimited escalation | Permission boundaries plus control of boundary removal and PassRole. |
| Allow a vendor to assume a customer role safely | Scoped trust and appropriate external-ID conditions. |
| S3 permits access but decryption fails | Inspect KMS key policy, identity permissions and grants. |
Traps
- An external ID is not a secret authentication credential.
- Allowing access in an SCP does not grant it.
- Network reachability does not imply authorisation.
07 · Detection Engineering and Incident Response
Memory hook: Prepare, detect, contain, preserve evidence, eradicate, recover and learn.
Must remember
- Define owners, severity, escalation, communication and legal/evidence requirements before an incident. Pre-authorise scoped response roles, prepare a forensic account and test runbooks through simulations. Isolation should contain the threat without needlessly destroying evidence.
- An organisation trail centralises supported CloudTrail activity. Select needed data events explicitly; management events alone do not record every object access. Protect log storage with restrictive policies, encryption, retention and integrity validation appropriate to the evidence requirement.
- GuardDuty, Inspector, Macie and Security Hub findings answer different questions. Route findings through EventBridge to a deduplicated workflow. Tune severity and suppression carefully; a finding is a lead requiring context, not automatic proof of compromise.
- For suspected credential theft, identify the affected principal and sessions, scope exposure, revoke/contain access using supported mechanisms, rotate compromised secrets and investigate persistence. Deleting one access key does not necessarily invalidate every previously issued temporary session.
- For compute compromise, isolate network access using prepared controls, preserve relevant disk snapshots/logs and collect volatile evidence when required. Terminating immediately can lose memory and local data. Use dedicated forensic tooling and documented chain of custody; do not run arbitrary suspect binaries.
- Validate containment against existing connections, not only new ones. Security-group connection tracking can allow tracked flows to continue after rule changes; a replacement restrictive group alone is not universal proof that a compromised host has lost every active path. Select documented network/session controls appropriate to the incident while preserving the access needed for evidence collection.
- For exposed data, contain access, determine affected objects/versions and callers, inspect encryption and key use, and preserve the access evidence. Recovery needs a clean trusted baseline and validation that the attacker's persistence is gone.
- Troubleshoot missing detections by checking event scope, Region, organisation membership, service enablement, delivery permissions, KMS policy, retention and event-rule filters. A disabled log pipeline can look deceptively quiet. Monitor the monitoring system.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Suspected compromised EC2 with valuable evidence | Isolate and preserve before destructive remediation. |
| No S3 object API records in a trail | Check data-event selection and delivery, not only management logging. |
| Same finding repeatedly triggers remediation | Use durable deduplication/idempotency and state checks. |
Traps
- Rotation alone may not invalidate previously issued sessions.
- An empty dashboard can mean broken telemetry.
- Automated containment can disrupt evidence collection unless planned.
08 · Multi-Account Delivery and Supply-Chain Controls
Memory hook: Promote the same verified artifact through increasingly sensitive environments.
Must remember
- Separate source, build, test and deployment responsibilities. A central pipeline can assume deployment roles in workload accounts; scope trust, artifact-bucket permissions and KMS key access. Cross-account artifact access often fails because only the S3 policy was updated.
- Pin dependencies, verify provenance, scan artifacts and record hashes/digests. Sign supported artifacts and enforce verification where required. Store secrets in managed stores and retrieve them at runtime/build time with scoped access; never expose them in logs or intermediate artifacts.
- CodeBuild buildspec phases, environment variables, reports and cache settings affect repeatability. Use ephemeral build environments and avoid giving untrusted pull-request code production credentials. Pipeline approvals should display the exact artifact and evidence being approved.
- Choose in-place, rolling, blue/green, canary or linear deployment using capacity, availability and rollback constraints. Health checks, pre/post hooks and CloudWatch alarms must observe real application outcomes. A deployment that never receives representative traffic cannot validate its behaviour.
- Database changes require expand/migrate/contract sequencing or another compatible strategy. Keep old and new application versions able to work during the transition; delay destructive cleanup until rollback requirements are satisfied.
- Use CloudFormation/SAM/CDK-generated templates with linting, policy checks, change sets and isolated integration tests. StackSets and account provisioning scale baselines; delegated roles and permission boundaries constrain the automation itself.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Same release in staging and production | Promote an immutable artifact, rather than rebuild. |
| Untrusted contributor submits a pull request | Run isolated validation without privileged deployment credentials. |
| Cross-account pipeline cannot decrypt artifacts | Check customer-managed KMS permissions and trust as well as S3. |
Traps
- A passing scanner does not prove a package is harmless.
- Manual approval is weak if reviewers cannot see what they approve.
- Rollback of code does not automatically roll back data.
09 · Event Automation and Resilience Engineering
Memory hook: An event is a trigger, not permission to repeat a destructive repair forever.
Must remember
- EventBridge patterns select events; schedules run on time; target roles permit actions. Design retries, dead-letter handling, deduplication and event-age limits. Replays can repeat side effects, so action handlers need idempotency and current-state checks.
- Step Functions Standard workflows support durable orchestration patterns; Express has different execution/history semantics. Choose retry/catch, timeout, heartbeat and callback behaviour explicitly. Distinguish a task failure from an entire workflow failure.
- SSM Automation runbooks can coordinate remediation and approvals. Use concurrency/error thresholds, scoped document parameters and a tested rollback. A Config noncompliance event should not trigger a repair loop fighting an application deployment.
- Reliability includes dependency timeouts, bulkheads, circuit breakers, queues and load shedding. Scale on a metric related to useful demand, such as backlog per worker, and account for warmup/cooldown and downstream limits. More workers can overload a database.
- Recovery drills need measured RTO/RPO, restored keys/secrets/configuration, dependency ordering and traffic failback. Chaos experiments should have a hypothesis, bounded scope, stop conditions and observability. Test AZ/dependency failures rather than merely stopping a random server.
- Track deployment frequency, lead time, change failure rate and recovery time alongside SLOs. Error budgets connect reliability evidence to release decisions. Post-incident reviews should improve controls and runbooks, not just add an alarm for the last symptom.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Replay archived business events | Idempotent consumers and explicit replay scope. |
| Repair thousands of resources | Rate-limited runbooks with failure thresholds and approvals. |
| Queue backlog grows while database saturates | Address downstream capacity and backpressure before adding unlimited workers. |
Traps
- Exactly-once-looking orchestration does not guarantee every external side effect occurs once.
- An alarm threshold is not a service-level objective by itself.
- Failover without a failback plan leaves recovery incomplete.