Reviewed 10 October 2026 · Solutions Architect Professional
Memory hook: Trace organizational boundaries, data ownership, failure behavior and the full bill.
Organisational complexity 26%; new solutions 29%; continuous improvement 25%; migration/modernisation 20%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — Architect network connectivity strategies
- Choose peering, TGW, PrivateLink, shared VPCs and hybrid connections according to routing scope, tenancy and overlap. TGW association selects ingress routing; propagation advertises routes; subnet tables still matter. Hybrid DNS uses inbound/outbound Resolver deliberately. Central inspection needs flow symmetry, resilient paths and no bypass.
1.2 — Prescribe security controls
- Separate identity grants, trust, ceilings, resource policies, keys and network access. Apply Organizations/OUs, SCPs, boundaries and supported resource controls without blocking required global-service dependencies. Third-party trust needs precise principal/external-ID conditions; centralized log access needs S3 and KMS authorization.
1.3 — Design reliable and resilient architectures
- Set RPO/RTO per business service and include identity, DNS, keys, secrets, quotas and artifacts. Separate AZ resilience from regional DR and compromise isolation. Test real failover/failback and data reconciliation; a healthy endpoint is not evidence of safe write ownership.
1.4 — Design a multi-account AWS environment
- Use account boundaries for security, lifecycle and cost ownership; landing zones establish identity, logging, networking and controls. Delegated administrators reduce management-account use. Baselines need repeatable provisioning, OU placement, emergency access and controlled exceptions, not just account creation.
1.5 — Determine cost optimization and visibility strategies
- Use account/tag ownership, activated allocation tags, Cost Explorer and detailed billing exports. Attribute shared network/platform costs. Rightsize first, then match commitments to measured stable usage; monitor utilization separately from coverage. Budgets/anomaly detection inform rather than universally cap spend.
2.1 — Design a deployment strategy to meet business requirements
- Match delivery to availability and rollback: rolling, immutable, blue/green, canary or linear deployment. Promote signed/identified artifacts through isolated environments. Database expand/migrate/contract preserves compatibility while old and new code coexist; code rollback alone cannot reverse destructive data changes.
2.2 — Design a solution to ensure business continuity
- Backup/restore minimizes idle expense but requires restoration time; pilot light preserves core capability; warm standby runs reduced capacity; active-active adds conflict/ownership complexity. Cross-Region replication has service-specific lag/consistency; validate quotas, dependencies and traffic failover under realistic load.
2.3 — Determine security controls based on requirements
- Derive controls from classification, residency, tenant boundaries and threat model. Encrypt transport/storage, constrain keys and protect audit evidence separately from workloads. CloudFront origin authorization and viewer authorization are distinct. Retention/compliance locks create non-bypassable lifecycle obligations.
2.4 — Design a strategy to meet reliability requirements
- Use fault isolation, bounded dependency timeouts, circuit breakers, queues and graceful degradation. Multi-AZ application capacity still depends on database, NAT/DNS and identity paths. Define health checks by business usefulness and ensure standby paths can carry recovery load.
2.5 — Design a solution to meet performance objectives
- Convert objectives into latency percentiles, throughput and freshness. Identify CPU/memory, IOPS/throughput, connections, query plans, hot keys and network limits. Cache/replicas reduce selected read demand; proxying addresses connections; partitioning changes query/operating complexity.
2.6 — Determine a cost optimization strategy to meet solution goals and objectives
- Compare complete solution cost: compute, licenses, storage minimums, requests, NAT/endpoints/transit, cross-AZ/Region movement, logs and staff effort. Managed/serverless options reduce some operations but can cost more for certain steady/high-volume patterns. Choose the cheapest design that meets all constraints.
3.1 — Determine a strategy to improve overall operational excellence
- Use SLOs, traces, metrics, runbooks, reviewed IaC, deployment evidence and incident learning. Automate safe repetitive work with bounded roles and retries. A green infrastructure check may miss business failure; monitor critical user transactions and telemetry delivery.
3.2 — Determine a strategy to improve security
- Inventory access/exposure, remove long-lived credentials, tighten boundaries/trust and separate key/log administration. Improve preventive, detective and responsive controls using actual findings. Prioritize impact and feasibility; test inherited policies and service dependencies before organization-wide enforcement.
3.3 — Determine a strategy to improve performance
- Profile the request before scaling. Query/index changes, caching, batching, connection reuse, partition distribution and storage tuning address different bottlenecks. Cold caches and stampedes can overload dependencies after failover; test recovery performance as well as steady-state speed.
3.4 — Determine a strategy to improve reliability
- Remove shared failure domains and validate recovery paths with fault/restore exercises. Replication preserves availability but can propagate corruption. Protect history and keys; reconcile writes before failback. Reduce toil and operational uncertainty through reproducible recovery automation.
3.5 — Identify opportunities for cost optimizations
- Analyze idle resources, rightsizing, data movement, unused replicas, log retention, object versions and commitment waste. Schedule eligible environments; lifecycle eligible data only after understanding retrieval/retention economics. Cheapest hourly instance is not always cheapest completed workload.
4.1 — Select existing workloads and processes for potential migration
- Inventory business services, dependencies, licensing, batch jobs, identity and hidden consumers. Assess value, readiness, risk and deadline before grouping migration waves. Discovery evidence and application-owner validation prevent moving one tightly coupled component without its dependencies.
4.2 — Determine the optimal migration approach for existing workloads
- Choose retain, retire, rehost, relocate, repurchase, replatform or refactor by required change and time. Application Migration Service supports server rehosting; DMS supports database data migration/CDC; DataSync transfers files/objects. Schema/application conversion is a separate compatibility exercise.
4.3 — Determine a new architecture for existing workloads
- Modernize selectively: externalize sessions, decouple with queues/events, use suitable managed data services and preserve required contracts. A strangler approach routes selected capabilities to new code while legacy remains. Define authoritative data ownership and anti-corruption boundaries to prevent uncontrolled dual writes.
4.4 — Determine opportunities for modernization and enhancements
- Sequence baseline load, CDC, reconciliation, controlled source writes, drained lag, client switch and business validation. Rollback must account for writes at the new target. Transactional outbox/sagas address cross-component consistency patterns but still need idempotency and valid compensation. Retire temporary access/resources after acceptance.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Many accounts need one private service, CIDRs overlap | PrivateLink when its service pattern fits. |
| One central firewall must inspect every flow | Symmetric routes and resilient appliance placement, with tested failure paths. |
| Short data-center exit deadline | Dependency-aware rehost/replatform may precede selective refactoring. |
| Regional failback after writes at the recovery site | Reconcile/synchronize data and establish writer ownership before traffic switch. |
| Cross-account encrypted snapshot inaccessible | Check sharing support and customer-managed key/data authorization. |
| Cost report has large shared-network spend | Use agreed allocation plus complete traffic-path analysis, not only workload instance tags. |
Traps
- Lowest operational effort is a deciding criterion only when the scenario asks for it.
- DNS switching cannot reconcile divergent writes.
- A backup connection too small for recovery load is not adequate resilience.
- SCPs do not govern management-account principals and service-linked roles like ordinary member principals.
- CloudFormation retention and data-retention controls can intentionally leave resources after stack deletion.
Verification cues
- Draw account/Region/AZ boundaries and label the identity, key and route required for every dependency.
- For each migration, state cutover conditions, validation evidence, accepted data-loss window and how rollback handles new writes.
- For an optimization, name the measured bottleneck, expected improvement and risk to reliability/freshness before recommending it.
Last-pass active recall
1. What is the difference between TGW association and propagation?
Association chooses the attachment’s ingress route table; propagation supplies its routes to selected tables.
2. Why is active-active hard for mutable data?
Concurrent writes need supported consistency/conflict handling and clear ownership; traffic routing alone does not solve this.
3. What does a successful DMS task not prove?
That every schema feature, stored procedure, query and business total is compatible/correct.
4. Why isolate audit storage from workload administrators?
A compromised workload should not be able to erase the evidence needed to investigate it.
5. What should happen before buying a large commitment?
Rightsize, remove waste and measure a stable eligible baseline and future change risk.
Sources and version check
- Official SAP-C02 exam guide
- AWS Well-Architected
- AWS migration strategies
- Disaster recovery strategies
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Getting Started with AWS
Memory hook: Choose the geography first, separate failure domains next, and bring eligible content closer to users at the edge.
Must remember
Region, AZ and edge answer different questions
- A Region is a geographic deployment boundary containing multiple Availability Zones. Regions are designed for isolation; launching something in one does not automatically create a copy in another.
- An Availability Zone (AZ) is one or more data centers with independent infrastructure and low-latency connectivity to other AZs in that Region. Deploying working replicas across AZs addresses a local infrastructure failure.
- AZ names such as us-east-1a can map differently between accounts; AZ IDs identify the same physical AZ across accounts. Shared networking designs should not assume matching letters mean matching locations.
- A VPC is regional; each subnet belongs to one AZ. Instances inherit the AZ of their subnet. EBS volumes and ENIs also have AZ boundaries, so they cannot simply move with an application across AZs.
- Edge locations support delivery and DNS services such as CloudFront and Route 53. They are not ordinary AZs where you launch an RDS database.
- Local Zones extend selected regional services closer to users; AWS Wavelength addresses supported telecom/5G-edge workloads; AWS Outposts brings AWS infrastructure to customer premises. These are different deployment models, not alternate names for CloudFront caches. Details belong in migration and recovery.
Choose a Region using requirements, not habit
- Apply hard constraints first: data residency, legal requirements, required service/feature availability and organizational restrictions. Check replication, backups and cached copies too.
- Then compare measured end-to-end latency, regional prices, operational support and recovery needs. The closest city is not proof of the best actual application latency.
- Multi-AZ usually addresses availability within a Region. Multi-Region addresses geographic recovery or global serving, adding replication, consistency, failover and cost decisions.
- A second subnet or a standby deployment description does not itself provide active application capacity. Ask what is running, where data resides, and how clients reconnect when a component fails.
Scope and responsibility
- IAM and Route 53 have global control planes. EC2, VPC and RDS use regional resources. S3's shared bucket-name namespace does not make the bucket's data globally replicated.
- A provider or CLI default Region selects an API context; it does not migrate existing resources or change every global service.
- The AWS Management Console is a browser interface to services; console access does not bypass IAM or automatically select the correct account and Region.
- Under shared responsibility, AWS secures underlying cloud facilities, hardware and infrastructure. Customers remain responsible for identities, data classification, access permissions and their service configuration.
- For normal EC2, the customer patches the guest OS and application. With managed databases, AWS manages more of the underlying platform, but customers still design access, schemas, backup choices and recovery procedures. More managed does not mean responsibility-free.
Choose under exam pressure
| Requirement or clue | Decision and reason |
|---|---|
| Survive one AZ failing | Working capacity and data resilience across AZs |
| Recover after a regional disaster | Cross-Region recovery plan with explicit RPO/RTO |
| Worldwide users download public images | Edge caching, if global copies are permitted |
| Customer records must remain in a jurisdiction | Approved regional storage, backup and replication locations |
| Same AZ across accounts matters | Compare AZ IDs, not only AZ letters |
| AWS compute required on premises | Consider Outposts rather than a CDN edge |
Traps
- Availability is not geographic recovery. Multiple AZs remain inside one Region; an edge cache is not a complete database recovery plan.
- A global name is not a global copy. Know the data location separately from the resource's naming or control-plane scope.
- A managed service still needs a customer design. Wrong access policies or missing application recovery can undermine reliable infrastructure.
02 · IAM & AWS CLI
Memory hook: Identify who is calling, how they authenticated, and which policies authorize the exact request.
Must remember
Identities and credentials
- The root user has special account-level capabilities. Protect it with MFA, avoid routine use, and do not create root access keys. Delegate ordinary administration; use root only when the specific operation requires it.
- An IAM user is a persistent identity. Console passwords sign in to the console; access keys sign programmatic requests. These are different credentials.
- An IAM group collects users and grants common permissions. Groups cannot contain other groups, are not shared logins and cannot be assumed like roles.
- A role is an assumable identity issuing temporary credentials through AWS STS. Those credentials include an access key ID, secret key and session token, and expire.
- Prefer federation/IAM Identity Center for workforce access and roles for applications. External SAML/OIDC identities can federate to appropriate AWS access; do not create a permanent IAM user for every workload.
- A role's trust policy controls who can assume it; its permission policies control what the assumed identity can do. Cross-account role access needs appropriate caller authorization and target trust.
- For EC2, an instance profile exposes the role to the instance. Lambda and other services use their own service-role arrangements; a role is not always an instance profile.
Read and evaluate a policy
- A JSON policy contains Version and Statement; statements use Effect, Action, Resource and optional Condition. Resource policies also identify Principal. Sid is an optional statement label.
- Version selects the policy language, not when the policy was last edited. Action identifies API capabilities; Resource identifies the affected ARN or wildcard where the action requires one.
- Requests are implicitly denied without an applicable authorization. Explicit deny overrides allow. Conditions such as requiring secure transport determine whether a statement matches.
- Identity policies and resource policies can participate together. Permissions boundaries, session policies and organization policies can restrict effective permissions; they do not make an otherwise unauthorized request universally allowed.
- See advanced IAM for principal/session, boundary, SCP and cross-account nuances; “all Allows add together” is unsafe.
Authentication hygiene and evidence
- MFA adds an authentication factor; it grants no permissions. A password policy controls IAM-user password requirements, not API key safety or federated identity-provider settings.
- Prefer short-lived credentials and least privilege. Never embed permanent keys in source code, AMIs, user data or Terraform variables.
- AWS CLI profiles select configuration and credentials. Environment variables or cached sessions can affect the actual caller; always distinguish intended profile from effective identity.
- CloudShell runs a managed shell using the signed-in console identity. It is not automatically an administrator and does not bypass IAM.
- Credential report: IAM-user credential status, password/key age and MFA information across the account.
- Access Advisor: historical service access that helps identify overbroad permissions. Access Analyzer has different policy/resource-access analysis capabilities. History alone cannot prove a permission will never be needed.
Choose under exam pressure
| Requirement or clue | Decision and reason |
|---|---|
| EC2 application must read one S3 prefix | Scoped instance role; no embedded keys |
| Employees need central multi-account login | Federation/Identity Center with temporary access |
| Team of IAM users needs the same permissions | IAM group with shared policies |
| Audit old or unused user credentials | Credential report |
| Reduce unused service permissions | Access Advisor plus workload requirements |
| Another account needs temporary access | Trusted role with appropriate permissions |
| HTTPS required despite an identity Allow | Matching explicit Deny blocks insecure requests |
Traps
- Authentication is not authorization. A valid login or API signature does not prove a requested action is allowed.
- A policy grants nothing merely by existing. The identity attachment or resource relationship must make it applicable.
- Display redaction is not secret protection. Local state and configuration can retain credential material even when the UI masks it.
03 · EC2 Fundamentals
Memory hook: Match the machine to the bottleneck, the purchase model to the commitment, and access to a temporary role.
Must remember
Compute and bootstrap
- An AMI supplies the launch image; the instance type supplies CPU, memory, networking and supported storage characteristics. Both must support the chosen CPU architecture.
- Choose compute optimized for CPU-heavy work, memory optimized for large in-memory datasets, storage optimized for demanding local I/O, and accelerated instances for supported GPU/accelerator workloads. General purpose balances resources.
- T-family burstable instances use CPU credits. Standard and unlimited credit modes have different sustained-load and billing consequences; a small headline hourly price is not the whole workload cost.
- Graviton/ARM instances require compatible OS images, binaries and containers. A lower-cost architecture is not a drop-in replacement for an incompatible binary.
- User data usually bootstraps on first launch and runs with elevated privileges. A stopped/started instance does not automatically rerun ordinary first-boot logic. Keep secrets out of it.
- A maintained golden AMI preinstalls software; user data finishes environment-specific configuration.
Connectivity and credentials
- Security groups are stateful allow rules attached to interfaces. Rules from associated groups combine; SGs do not have explicit deny rules. Return traffic for an allowed flow is tracked.
- Restrict sources using CIDRs or supported SG references; references identify source interfaces, not inherited rules.
- Common ports: SSH 22, FTP control 21, HTTP 80, HTTPS 443, RDP 3389, DNS 53, SMTP 25/587, NFS 2049. FTP data behavior and DNS TCP/UDP needs depend on the protocol use case.
- SSH requires the right OS user/key and a reachable network path. EC2 Instance Connect supplies temporary SSH public-key access but still needs connectivity; an Instance Connect Endpoint is a separate networking option.
- Session Manager needs an agent, IAM and service-endpoint access, without inbound SSH. Private instances can use VPC endpoints.
- An instance role/profile provides temporary AWS credentials; IMDSv2 requires session tokens. Local software should use the role credential chain instead of permanent keys.
Buying capacity and controlling cost
| Workload clue | Choice | Main tradeoff |
|---|---|---|
| Short or unpredictable use | On-Demand | No long-term discount commitment |
| Stable matching configuration | Reserved Instances | One/three-year commitment; flexibility varies |
| Stable dollar-per-hour usage | Savings Plans | Spend commitment; flexibility depends on plan type |
| Interruptible, restartable processing | Spot | Spare capacity can disappear |
| Physical licensing/placement control | Dedicated Host | Host management and licensing economics |
| Single-tenant hardware | Dedicated Instances | No equivalent physical host placement control |
| Need matching capacity in a particular AZ | Capacity Reservation | Unused reserved capacity can still bill |
- Standard versus Convertible RIs: discount/flexibility differ; Convertible provides supported exchanges. Regional RIs provide a billing benefit without a capacity reservation; zonal RIs reserve matching AZ capacity. Savings Plans do not themselves reserve capacity.
- Spot Fleet/EC2 Fleet can diversify instance types and AZ capacity pools. Use retries, checkpoints and interruption-tolerant design; do not rely on one irreplaceable stateful worker.
- Spot stop/terminate notices normally provide two minutes, on a best-effort basis. Hibernation interruption begins immediately rather than providing that same advance window.
- AWS Budgets alerts use delayed billing data and need sufficient history for forecasts. An alert is not a hard real-time cap. EBS, snapshots, public IPv4 and commitments can keep billing after an instance stops.
Choose under exam pressure
| Requirement | Decision and reason |
|---|---|
| CPU saturated while RAM remains free | Investigate compute sizing before buying more memory |
| Batch jobs can retry on another worker | Spot with resilient orchestration and diversification |
| Must launch in one AZ at a known time | Capacity guarantee, not only a discount |
| Administer private instances without SSH keys | Session Manager with endpoint access and roles |
| Application must access another AWS service | Least-privilege instance role |
Traps
- Discount and guaranteed capacity differ. A Savings Plan cannot fix insufficient capacity in a chosen AZ.
- Stopped is not deleted. Retained storage, IP allocations and commitments may remain billable.
- Timeout and refusal suggest different investigations. Routing/SG/NACL issues often time out; no application listener may refuse the connection.
04 · EC2 SAA Level
Memory hook: Preserve the right identity, separate the right failure domains, and distinguish resuming a machine from recovering an application.
Must remember
Addresses and interfaces
- Private IPv4 addresses belong to the VPC address space. An instance's primary private IP normally remains with its primary ENI through stop/start; terminating resources is a different lifecycle operation.
- An automatically assigned public IPv4 is not a durable endpoint and can change on stop/start. Reboot and stop/start are not equivalent operations.
- An Elastic IP (EIP) is a static regional IPv4 allocation that can be reassociated with supported resources. A static address simplifies allowlists or appliance failover, but allocation itself supplies no redundant application capacity.
- Public IPv4 addresses incur charges, including allocated Elastic IPs kept idle. Prefer an appropriate load-balancer/DNS endpoint when the real requirement is replaceable application servers rather than a single machine's address.
- An elastic network interface (ENI) contains private IPs, a MAC address and security-group associations. It belongs to a subnet and therefore to one AZ.
- A secondary ENI can move between compatible instances in the same AZ, allowing an appliance to retain a private network identity. The primary ENI cannot simply be detached from a running instance.
- Network-identity failover still needs detection, reassociation permissions, a ready replacement and application recovery. Moving an interface does not copy files, RAM or transactions to another machine.
Placement strategies
- Cluster placement packs supported instances close together in one AZ for tightly coupled, low-latency, high-throughput communication. Think HPC, not geographic disaster tolerance.
- Spread placement separates a small critical set across distinct underlying hardware. It reduces correlated hardware failure; it is not a replication or backup service.
- Partition placement separates groups of instances onto different underlying hardware. Partition-aware distributed applications such as Kafka, Cassandra or Hadoop can distribute replicas across those failure domains.
- Choose based on the application's communication pattern and failure model. A distributed data platform should know where it places replicas; separation is useful only if the application uses it.
- Placement/type support differs. Empty groups supply no running capacity, and this lab's burstable instance cannot demonstrate every strategy.
Hibernation and lifecycle
- Hibernation saves RAM to the encrypted EBS root volume, stops the instance, and restores the memory image on resumption. It suits expensive initialization held in memory.
- It requires supported instance/AMI configurations, prior enablement and adequate root-volume space. It is not interchangeable with an ordinary stop, which does not preserve RAM.
- Retained EBS storage continues to cost money. A stopped/hibernated machine is not an active standby serving users.
- Reboot restarts the OS. Stop ends compute while retaining eligible EBS state. Terminate destroys the instance, subject to volume-retention settings.
- Important instance-store data is not protected by hibernation or a durable-looking IP address. Storage lifecycle and replicas remain separate design decisions; review disks and shared files.
Choose under exam pressure
| Requirement or clue | Decision and reason |
|---|---|
| Low inter-node latency in tightly coupled HPC | Cluster placement on supported instances |
| Separate a small set of critical replicas | Spread placement plus application replication |
| Isolate groups in a distributed data platform | Partition placement and aware replica assignment |
| Retain a private appliance identity within one AZ | Secondary ENI failover |
| Fixed regional public IPv4 required | EIP, with explicit failover and cost consideration |
| Resume a long in-memory initialization | Supported hibernation |
| Continue operating through an AZ loss | Cross-AZ application/data design, not ENI movement |
Traps
- An ENI cannot be moved across AZs. Its subnet association is part of its placement boundary.
- Identity is not state. An IP address, hostname or MAC does not preserve an application's data.
- Suspension is not availability. Hibernation preserves working memory but does not serve requests while stopped or create a second failure domain.
05 · EC2 Instance Storage
Memory hook: Choose block, file or local scratch first, then check its failure boundary, performance and recovery behavior.
Must remember
EBS and instance store
- EBS is persistent block storage in one AZ. The attached EC2 instance must be in the same AZ. It can outlive an instance, but delete-on-termination settings determine each volume's lifecycle.
- Root/data volumes can have different deletion behavior. Stopped instances and unattached EBS volumes can retain storage charges.
- gp3 separates capacity from configurable IOPS/throughput and suits general workloads. gp2 performance depends on size and burst behavior. More GiB is not always the best way to solve an I/O bottleneck.
- io1/io2 Provisioned IOPS SSD suits demanding transactional random I/O and predictable performance. st1 and sc1 HDD target sequential throughput and colder sequential data; they are not boot-volume choices.
- Volume and instance EBS limits both constrain performance; more provisioned IOPS cannot bypass instance bandwidth.
- Multi-Attach is supported for eligible io1/io2 volumes and compatible Nitro instances in the same AZ, with engine/OS/Region restrictions. It is not supported for gp3 or as a boot-volume shortcut.
- Concurrent block writes require a cluster-aware filesystem/application and correct coordination/fencing. Ordinary filesystems mounted read-write from unrelated servers can corrupt data.
- Instance store is host-local and ephemeral on host loss, stop or termination, unlike reboot. Use it for reconstructable cache, scratch or replicated data, never the only durable copy.
Snapshots, images and encryption
- Standard EBS snapshots store incremental changed blocks. AWS manages dependencies, preserving later snapshots' restorability when older snapshots are deleted.
- Restore a snapshot into a new volume to change AZ; copy snapshots for cross-Region recovery.
- An EBS-backed AMI adds boot/launch metadata and references image snapshots. Deregistering the AMI and deleting snapshots are distinct cleanup decisions.
- Snapshot-created volumes can require initialization for full performance. Fast Snapshot Restore removes that initialization penalty for enabled snapshot/AZ combinations, with additional charges. It does not make the data newer.
- Snapshot Archive converts an incremental snapshot into a full archived snapshot. It trades lower long-term storage rates for retrieval delay, retrieval costs and a minimum storage duration; it is poor short-lived lab storage.
- Recycle Bin retains eligible deleted snapshots/AMIs for recovery, delaying final deletion and potentially retaining charges.
- KMS encryption covers EBS data and associated snapshots. Copying an unencrypted snapshot into an encrypted one supports migration; an existing unencrypted volume does not simply gain in-place encryption.
- Sharing an encrypted snapshot requires appropriate snapshot permissions and access to a suitable customer-managed KMS key. AWS-managed default keys cannot be shared as though they were your own cross-account keys.
EFS: files instead of blocks
- EFS is managed NFS for concurrent Linux filesystem access. Regional EFS distributes storage across AZs; One Zone has a different failure boundary.
- Mount targets provide network access in VPC subnets. NFS reachability, security groups, file permissions and any configured IAM access controls all matter.
- Elastic throughput adapts to demand; Provisioned throughput sets a chosen throughput level; Bursting links available throughput/credits to the storage workload. These are distinct from storage classes.
- Standard, IA and Archive classify data by access/storage economics. Lifecycle policies can move cold files; retrieval and minimum-duration charges can outweigh savings for short retention.
- EFS fits shared uploads/home directories, not every block-database workload; storage integrations covers specialized Windows/HPC filesystems.
Choose under exam pressure
| Requirement or clue | Decision and reason |
|---|---|
| Persistent instance boot disk | Supported SSD EBS volume |
| Predictable high random transactional I/O | Provisioned IOPS SSD plus adequate instance capability |
| Large sequential scans at lower cost | Appropriate HDD EBS type |
| Linux shared uploads across AZs | Regional EFS |
| Rebuildable high-performance scratch | Instance store |
| Repeatable bootable server image | AMI |
| Recover a disk in another AZ | Snapshot restore to a new volume |
| Long-retained rarely restored backup | Evaluate archive retrieval and minimums |
Traps
- Multi-Attach is not managed shared NFS. Shared blocks still require write coordination and remain in one AZ.
- A copy is only as current as its recovery point. More snapshots or faster initialization does not inherently improve the latest available RPO.
- Deleting compute is incomplete cleanup. Volumes, snapshots, AMIs, retention policies and filesystem data each have separate lifecycles.
06 · ELB & Auto Scaling
Memory hook: A load balancer chooses a destination; an Auto Scaling group maintains capacity; application state must survive either decision.
Must remember
Pick the right traffic layer
- Application Load Balancer (ALB) understands HTTP/HTTPS at layer 7. A listener receives traffic; ordered rules choose forward, redirect, authentication or fixed-response actions. Conditions can include host and path.
- Target groups hold instances, IPs or supported Lambda destinations, with health checks/attributes. Path rules can select different groups.
- Lower numbered listener priorities run first. A fixed response comes from the ALB and can succeed even when the application is unavailable.
- Network Load Balancer (NLB) handles TCP/UDP/TLS at layer 4 and supplies static addresses per enabled AZ, with optional Elastic IPs for supported internet-facing deployments. It fits non-HTTP traffic and IP allowlist requirements.
- NLB is not an HTTP path router. Modern NLBs support security groups in supported configurations; the old claim that NLBs never have security groups is unsafe.
- Gateway Load Balancer (GWLB) steers traffic through network appliances using GENEVE/UDP 6081, rather than routing application URLs.
- An internet-facing load balancer can forward to private backends; public users do not require public backend IPv4.
Connection behavior and health
- Stickiness uses supported cookies/affinity mechanisms to favor a target. It does not replicate session memory or guarantee that target will survive.
- Cross-zone load balancing allows a node to route across enabled AZs. ALB has it enabled at the load-balancer level, with target-group-level controls; NLB/GWLB default differently. Check transfer cost rules for the specific type.
- Deregistration delay drains in-flight requests during removal. Too little can interrupt long work; excessive values can slow scale-in and deployments.
- TLS termination uses listener certificates, often from ACM. SNI lets a compatible client indicate its hostname so a listener chooses the correct certificate. ALB certificates are regional; CloudFront ACM certificates use us-east-1.
- Target health checks detect application reachability on the configured port/path. EC2 status health and application health answer different questions.
- Health routing is not authorization: all-unhealthy/fail-open behavior can send traffic to unhealthy targets.
Scaling and replacement
- Vertical scaling changes machine size and may require interruption; horizontal scaling changes the number of workers. Externalized state makes horizontal replacement safer.
- AWS Auto Scaling plans coordinate resources. EC2 Auto Scaling manages instance groups; Application Auto Scaling handles supported dimensions such as ECS task counts. Plans can migrate to direct policies.
- Launch templates describe AMIs, type, user data, interfaces and related launch settings. Updating a template does not automatically update every already-running instance.
- An ASG maintains desired capacity between minimum and maximum bounds. Enable appropriate ELB health when application failures should cause replacement; EC2-only checks may miss a broken web process.
- Instance refresh rolls out launch changes with health, warmup and capacity constraints.
- Target tracking maintains a chosen metric target; step scaling changes capacity according to alarm severity; scheduled scaling anticipates known times; predictive scaling forecasts recurring demand.
- Choose a demand-related metric: CPU can fit compute-bound servers, ALB request count per target can fit web workers, and queue backlog per worker can fit asynchronous processing.
- Warmup/cooldown reduce unstable decisions during startup; grace periods do not prove readiness.
- Multi-AZ subnets with desired capacity one do not provide two active replicas. Production resilience needs sufficient surviving capacity and dependencies.
Choose under exam pressure
| Requirement | Decision and reason |
|---|---|
| Several web apps under paths or hostnames | ALB listener rules and target groups |
| Static addresses for TCP/UDP clients | NLB |
| Third-party network inspection fleet | GWLB |
| Scale before a known daily opening | Scheduled scaling |
| Maintain utilization near a target | Target tracking |
| Replace instances with a broken web process | ASG using relevant ELB health |
| Roll out a new AMI to existing capacity | Controlled instance refresh |
Traps
- Scaling and healing differ. Replacing an unhealthy instance can preserve the same desired capacity without adding demand capacity.
- A cookie is not a session database. Target loss still destroys target-local state.
- A successful listener test can bypass dependencies. Fixed responses do not prove backend, cache or database health.
07 · RDS, Aurora & ElastiCache
Memory hook: Replicas add read capacity, failover preserves service, backups recover history, proxies manage connections, and caches avoid repeated work.
Must remember
RDS: identify the actual bottleneck
- RDS provides managed relational engines. AWS operates much of the platform; you choose sizing, access and retention. Ordinary RDS provides no general guest-OS administration.
- RDS Custom allows supported OS/database customization within automation boundaries. Verify engine/platform fit; use EC2 when full control is essential.
- CPU/memory, storage, IOPS/throughput and connections are different bottlenecks; more disk does not fix CPU-heavy queries.
- Storage autoscaling increases storage toward a configured maximum; it does not shrink it. Monitor both consumption and the cap rather than treating it as unlimited.
- A DB subnet group describes eligible subnets; a parameter group configures engine settings. Neither creates running capacity or HA.
Read scaling, availability and recovery differ
- A classic Multi-AZ DB-instance deployment synchronously replicates to a standby for automatic failover. That standby does not serve application reads.
- A Multi-AZ DB cluster has a different architecture with readable standby instances. Always identify the deployment type before applying the shortcut “Multi-AZ is not for reads.”
- Read replicas principally offload reads, generally using asynchronous replication. Lag matters for read-after-write requirements. Promotion/cutover is a separate decision from the usual read-scaling purpose.
- Automatic failover changes the serving database; clients still need reconnection/retries and handling for interrupted transactions.
- Automated backups and transaction logs support point-in-time recovery (PITR). Restore produces a new database resource requiring a cutover; it does not undo selected changes in the running database.
- Manual snapshots persist until deleted and can support copying/sharing workflows. Encryption/key access and regional transfer/storage charges remain relevant.
- Replicas may reproduce accidental deletes or corrupt application writes. Replication is not backup history. RPO asks how much data loss is acceptable; RTO asks how long recovery may take.
Aurora: storage, endpoints and capacity
- Aurora separates compute instances from distributed cluster storage spanning multiple AZs. A one-instance Aurora cluster still has distributed storage, but adding an appropriately placed reader improves compute failover options.
- The writer/cluster endpoint follows the current writer. The reader endpoint distributes new connections among available readers; it does not redistribute every query inside one existing connection.
- Instance endpoints target individual instances. Custom endpoints select instance subsets, useful when reporting should use a different capacity group from ordinary application reads.
- Replica Auto Scaling changes the number of readers. Aurora Serverless v2 adjusts compute capacity. Neither should be confused with scaling the writer's count.
- Serverless v2 auto-pause/zero-capacity requires eligible versions/configuration; idle cost is not universally zero. Serverless v1 is retired; see service availability.
- Aurora Global Database uses cross-Region replication for global reads and recovery. It is a regional-disaster design, distinct from adding another reader in the same Region; asynchronous replication can imply unreplicated-write risk.
- Aurora cloning uses copy-on-write storage sharing for fast development/test copies. Shared unchanged pages and divergent writes differ from independent regional disaster-recovery copies.
- Babelfish for Aurora PostgreSQL supports many SQL Server interfaces to reduce application migration changes, but compatibility assessment is essential.
- Aurora ML integrations expose supported ML capabilities from database workflows; they do not turn every database engine into a model-training service.
Secure connections and protect capacity
- KMS protects data at rest, TLS protects data in transit, and supported IAM database authentication provides temporary authentication tokens. Database users/grants still authorize SQL operations.
- Security groups control network reachability. Permit the application SG on the required engine port; a private subnet alone is not a complete authorization design.
- Ports: PostgreSQL/Aurora PostgreSQL 5432, MySQL/MariaDB/Aurora MySQL 3306, SQL Server 1433, Oracle 1521, Redis OSS/Valkey 6379, Memcached 11211.
- RDS Proxy pools/reuses database connections and can help manage failover and connection surges. It is not a query-result cache and does not remove all database capacity limits.
- IAM authentication and Secrets Manager solve different integration needs; see encryption and secrets.
ElastiCache and cache correctness
- Redis OSS/Valkey offer richer structures such as sorted sets, with replication, failover and persistence capabilities depending on deployment.
- Memcached is a simpler multithreaded distributed cache. Do not assume every modern Serverless feature matches a classic node-based engine comparison; choose the actual deployment's capabilities.
- Lazy loading/cache-aside: read cache, fetch the database on a miss, then populate. It avoids caching never-read objects but adds miss latency and can return stale data.
- Write-through: update the cache alongside database writes. It improves hit availability at extra write work, but still requires a coherent failure/invalidation strategy.
- TTL expires entries. Stagger expiry or use appropriate request coordination to avoid a cache stampede when many entries expire together.
- An external session store makes web servers replaceable. Required session durability and failover must still be designed; stickiness alone does not preserve lost memory.
- Cache authentication/user controls, supported IAM authentication, TLS, and SGs address distinct layers. A cache is not safe merely because it has no public endpoint.
Choose under exam pressure
| Requirement or symptom | Decision and reason |
|---|---|
| Primary overloaded by tolerant reporting reads | Read replicas; account for lag |
| Automatic recovery from AZ failure | Appropriate Multi-AZ deployment |
| Recover before an accidental update | PITR or suitable snapshot |
| Burst of short-lived Lambda connections | RDS Proxy |
| Repeated expensive reads or fast shared sessions | Appropriate cache strategy |
| Isolate Aurora analytics readers | Custom endpoint and selected capacity |
| Variable Aurora compute demand | Serverless v2, with supported limits/cost model |
| Database recovery after regional loss | Cross-Region copies or Global Database design |
Traps
- “Multi-AZ” is not one architecture. Distinguish classic DB-instance standbys from readable DB-cluster instances.
- Encryption is not one switch. At-rest KMS, network TLS, identity authentication and SQL privileges solve different problems.
- Caches and replicas copy failures too. Neither automatically supplies a clean historical recovery point.
08 · Route 53
Memory hook: DNS chooses an answer that clients may cache; it does not inspect or balance every application request.
Must remember
Resolution, records and ownership
- A recursive resolver finds an answer on a client's behalf, following cached information and authoritative DNS as needed. An authoritative hosted zone contains the records for its namespace.
- A maps a name to IPv4; AAAA to IPv6; CNAME to another hostname; NS identifies authoritative name servers; MX identifies mail servers; TXT carries text/verification data; SOA carries zone metadata.
- TTL controls how long a DNS answer may be cached. Lower TTL can improve the responsiveness of future changes but increases queries and cannot instantly invalidate previously cached answers.
- Domain registration and DNS hosting are separate services. A third-party registrar can delegate a domain to Route 53 by using the correct Route 53 name servers. Moving DNS does not necessarily require transferring registration.
- Creating a same-named public hosted zone does not configure delegation; resolvers need the correct authoritative chain.
- A CNAME cannot occupy a zone apex alongside its required SOA/NS records. Route 53 Alias A/AAAA can map an apex to supported targets such as an ALB or CloudFront distribution.
- Alias is an AWS DNS feature with supported target rules, not permission to point any apex record at an arbitrary hostname. Its TTL/health behavior depends on the target.
Every routing policy answers a different question
| Requirement or clue | Routing policy and meaning |
|---|---|
| Ordinary answer for one service | Simple: basic response without specialized selection |
| Gradual rollout or relative distribution | Weighted: choose among records according to relative weights |
| Best response latency among configured Regions | Latency: use measured latency information, not just map distance |
| Primary/standby DNS recovery | Failover: prefer healthy primary, otherwise secondary |
| Country/continent or location-specific content | Geolocation: select by user location; define a default |
| Shift a geographic catchment boundary | Geoproximity: resource/user location with adjustable bias |
| Known client-source CIDR requirements | IP-based: use CIDR collections to choose destinations |
| Several healthy IP answers | Multivalue: return a small set of healthy answers; not a full load balancer |
- Weights are relative, not required to total 100. Resolver caching and client reuse mean a 90/10 policy does not guarantee nine of each ten HTTP requests use one endpoint.
- Latency and geography differ: the geographically nearest endpoint need not have the lowest measured network latency.
- Geolocation is not authorization or residency enforcement; storage locations, cache distribution and access controls need separate policies.
- Geoproximity bias changes the area attracted to a resource; it is not the same as changing an exact percentage weight.
- Traffic Flow can compose visual traffic policies, with separate pricing/management considerations. It is a configuration facility rather than another universal per-request proxy.
Health and failover
- An endpoint health check probes a supported publicly reachable endpoint using its configured protocol and conditions.
- A calculated health check combines child checks using a threshold; a CloudWatch alarm-based health check derives health from supported alarm/metric behavior instead of a direct public probe.
- Route 53 public health checkers cannot directly reach an ordinary private-only IP. A suitable metric/alarm integration is one way to express private resource health.
- Health checks are separate resources and must be correctly associated with the DNS design. Supported Alias targets may provide Evaluate Target Health behavior instead of needing a duplicate direct endpoint probe.
- Failover depends on detection time, DNS answers, resolver caches and application reconnection. A small TTL is not a promise that all clients switch instantly.
- Secondary endpoints still need usable capacity and sufficiently current data; health checks preserve neither transactions nor state.
Private zones and hybrid DNS
- Private hosted zones answer within associated VPCs through the appropriate resolver context. Required VPC DNS attributes, associations and application resolver configuration must be correct.
- Split-view DNS uses the same namespace with different internal and external answers. A private zone can intentionally shadow public names.
- If an associated matching private zone lacks the requested name/type, resolution can return NXDOMAIN, rather than automatically falling back to the public zone.
- Route 53 VPC Resolver is the current name for the VPC service historically called Route 53 Resolver. Its inbound endpoints receive queries from on-premises/other connected networks.
- Outbound endpoints and forwarding rules send matching VPC queries to external DNS servers. Think “inbound to the VPC” and “outbound from the VPC.”
- Endpoints do not create the underlying VPN/Direct Connect route. DNS ports, security groups, routes, forwarding rules and resilient endpoint placement are separate requirements.
- Private zones do not support every public-zone routing policy.
Choose under exam pressure
| Situation | Decision and reason |
|---|---|
| Apex domain must reach an eligible AWS load balancer | Alias A/AAAA, not apex CNAME |
| Keep registrar but use Route 53 DNS | Update authoritative delegation |
| Hybrid clients must resolve AWS private names | Inbound Resolver endpoint and private connectivity |
| VPC clients must resolve corporate names | Outbound endpoint with matching forwarding rules |
| Internal name exists publicly but fails inside VPC | Inspect private-zone shadowing and record/type |
| Exact request-level canary split required | DNS weighting alone cannot guarantee it |
| Private backend needs failover health | Appropriate alarm/metric-based health integration |
Traps
- Resolution is not connectivity. A correct address does not open a firewall or establish a route.
- Caching limits immediate control. DNS updates do not terminate established connections or flush every resolver.
- Private and public evidence differ. A laptop using public DNS cannot prove a VPC-only record is absent.
09 · Classic Solutions Architectures
Memory hook: Make compute replaceable by putting necessary state in services with deliberately chosen durability, access and failure behavior.
Must remember
Case 1: a public status service
- A service generating independent responses can begin with one EC2 instance/EIP, but one instance remains one serving failure point.
- Route 53 supplies stable naming; an ALB distributes HTTP requests; an ASG maintains/scales the fleet. Naming, routing and capacity lifecycle are separate responsibilities.
- Run sufficient healthy capacity across AZs. Two eligible subnets with desired capacity one offer placement choices, not two active application replicas.
- An EIP preserves a regional public address, but reassociation is not equivalent to distributing traffic across a resilient fleet.
- Keep the web tier stateless: any healthy instance should handle the next request. Necessary state can remain in other tiers.
Case 2: a shopping-session service
- A cart in one server's RAM depends on that server. ALB stickiness provides affinity; failure or scale-in can still destroy local memory.
- Move sessions to an appropriate ElastiCache/Valkey/Redis or other shared store. Choose replication, failover, persistence and expiry according to the consequences of losing sessions.
- Cache repeated product reads to reduce database load, with invalidation/TTL. Carts, inventory and descriptions can have different freshness requirements.
- Read replicas offload suitable reads, not general writes. Replication lag matters immediately after purchases or inventory changes.
- Multi-AZ databases address availability; backups preserve recovery history. Database choices distinguishes readable clusters from classic non-readable standbys.
- Security-group chaining: clients→ALB, ALB SG→application SG, application SG→database/cache SG on required ports. SG references identify sources; they neither copy rules nor create routes.
Case 3: a collaborative editorial site
- EFS fits shared uploads when Linux applications expect filesystem semantics. Use appropriate regional storage/access resilience.
- Aurora or another relational database stores structured content. Filesystems are not databases, and caches do not replace persistent uploads.
- If applications support object APIs directly, S3 can serve media through an appropriate access/CDN design. It is not interchangeable with shared POSIX operations without integration.
- Scale and secure each tier independently. Web availability depends on filesystem, database, session-store and network availability too.
- Managed services still need backups, deployment rollback and recovery testing.
Faster launches and safer replacement
- A golden AMI preinstalls common dependencies, shortening launch time but requiring image maintenance and patching.
- User data finishes environment-specific bootstrap. Long installs and external dependencies delay readiness; repeatable scripts reduce replacement surprises.
- Snapshot restoration recovers disk content. Application consistency, initialization performance and endpoint cutover determine usable recovery time.
- Meaningful ASG/ELB health should check application readiness, not merely a running OS.
- Deregistration delay allows in-flight work to finish; it does not make target-local state durable.
Managed application environments
- Elastic Beanstalk deploys supported applications using underlying AWS resources such as EC2, load balancers and scaling groups. Configuration choices and resource charges remain.
- An application organizes versions/environments; an environment is a running deployment; a supported platform supplies runtime/web-stack configuration.
- Web-serving environments handle incoming application traffic. Worker environments can process queued background work, separating synchronous responses from slower tasks.
- Deployment strategies trade interruption, rollback capability and temporary extra capacity. Managed deployment is not inherently serverless billing.
- Keep durable data independent of disposable environments when recreating application infrastructure must not delete business records.
Choose under exam pressure
| Requirement or symptom | Architectural response |
|---|---|
| Carts vanish during scale-in | Externalize sessions with suitable durability |
| Shared Linux uploads | EFS with resilient access/storage |
| Repeated reads overload DB | Cache and/or suitable read replicas |
| Slow installation delays scaling | Golden AMI plus short bootstrap |
| Process dies but EC2 stays healthy | ELB health integrated with ASG replacement |
| Supported web deployment with less infrastructure operation | Elastic Beanstalk |
| Recreate environment without losing data | Independent persistent-data lifecycle |
Traps
- Stateless compute is not a stateless system. External stores become dependencies needing their own resilience.
- Managed is not automatically cheap or highly available. Capacity, configuration and retained data still matter.
- Solve the failed requirement. Stickiness cannot provide durable sessions; more read replicas cannot resolve write contention.
10 · S3 Introduction
Memory hook: A key names an object, versioning preserves history, and replication places eligible copies elsewhere.
Must remember
Objects, access and current limits
- S3 stores objects: data plus metadata identified by bucket and full key. It is not an EBS block device or NFS filesystem.
- General-purpose buckets use flat keys; slash-separated prefixes resemble folders but also support policy/lifecycle selection.
- Names may use the shared global namespace within a partition or current account regional namespaces. The bucket's chosen Region determines data location; global naming does not imply global replication.
- Strong read-after-write consistency covers object writes/deletes and corresponding reads/listing. Asynchronous replication and CDN caches still have independent delay.
- Current limits, checked 2026-10-09: AWS's upload guide states a 50 TB multipart maximum; its specification expresses the bound as 48.8 TiB. Older 5 TB course limits are outdated. A single PUT remains 5 GB.
- Multipart supports 10,000 parts, each 5 MiB–5 GiB, except no minimum for the final part. Choose adequate part sizes; unfinished parts bill until completed/aborted. See S3 operations.
Policies and website delivery
- IAM identity policies authorize callers; bucket policies attach rules to resources. Bucket actions need bucket ARNs; object actions need appropriate object ARNs.
- Explicit deny wins. Block Public Access restricts public exposure but does not grant private application access.
- Bucket owner enforced Object Ownership disables ACLs and simplifies ownership; policies handle access.
- Static websites serve HTML/JavaScript/assets, not server-side PHP or similar code. S3 website endpoints do not directly provide HTTPS.
- Private-origin HTTPS delivery commonly uses CloudFront with S3 REST-origin/OAC. Website endpoints are custom origins with different access behavior; see edge delivery and object security.
Versions and replication
- Versioning preserves prior data when a key is overwritten. Version ID, object key and bucket name identify different things.
- A normal delete usually creates a delete marker; earlier versions remain billable. Deleting a specific version permanently removes that version.
- Suspending versioning does not erase earlier history. Manage noncurrent-version retention separately from current objects.
- CRR copies across Regions; SRR stays in one Region. Both need versioning, rules and appropriate permissions.
- Live replication is asynchronous and normally covers eligible writes after configuration. Existing objects need explicit backfill, such as Batch Replication.
- Delete-marker and encrypted-object replication require suitable rule/key configuration. Never assume every deletion or protection setting propagates identically.
- A replica is not historical backup by itself; a CDN cache is not durable regional replication.
Storage classes: access, failure boundary and cost
| Requirement | Class and tradeoff |
|---|---|
| Frequent access and regional resilience | Standard |
| Unknown/changing access | Intelligent-Tiering, with monitoring/tiering economics |
| Infrequent, immediate access | Standard-IA, with retrieval/minimum charges |
| Re-creatable infrequent data; one AZ acceptable | One Zone-IA |
| Archive requiring immediate access | Glacier Instant Retrieval |
| Archive tolerating restore delay | Glacier Flexible Retrieval |
| Long archive tolerating longer recovery | Glacier Deep Archive |
| Low-latency object access colocated in an AZ | Express One Zone, using directory buckets |
- Durability means avoiding data loss; availability means successful access now. A durability figure is not an uptime percentage.
- One Zone classes accept a different failure boundary from regional classes. Avoid storing the only irreplaceable copy there when AZ-loss survival is required.
- Intelligent-Tiering adapts eligible access tiers; optional archive tiers change retrieval behavior. It is not guaranteed cheaper for every object.
- Retrieval, requests, minimum duration/size and transfer charges can outweigh lower storage rates. Short-lived tiny objects are poor archival candidates.
- Directory buckets/Express One Zone have different feature support from general-purpose buckets; do not assume identical versioning or replication behavior.
Choose under exam pressure
| Requirement | Decision and reason |
|---|---|
| Recover overwritten data | Versioning with suitable retention |
| Durable copy of new writes in another Region | CRR with permissions/monitoring |
| Replicate older objects too | Explicit backfill |
| Private static content globally over HTTPS | CloudFront and private S3 REST origin |
| Only copy must survive AZ loss | Suitable regional storage |
| Archived objects need immediate reads | Glacier Instant Retrieval |
Traps
- Invisible is not deleted. Versions, markers and multipart parts have independent retention.
- Strong consistency does not control caches or replication.
- Storage price is not total workload price. Access and retention charges can reverse apparent savings.
11 · Advanced S3
Memory hook: Measure access, choose an economical storage class, automate aging, and make event consumers tolerate retries.
Must remember
Storage cost is more than the monthly rate
- Standard suits frequent access and has no minimum storage duration. Intelligent-Tiering responds to changing access patterns; monitoring charges and small-object eligibility matter. Its optional archive tiers require restore workflows.
- Standard-IA and One Zone-IA: 30-day minimum storage duration; retrieval charges; 128 KB minimum billable object size. One Zone-IA fits recreatable data because it does not survive loss of its AZ.
- Glacier Instant Retrieval: immediate access, but a 90-day minimum. Glacier Flexible Retrieval: restore before reading, with minutes-to-hours retrieval choices and a 90-day minimum. Deep Archive: hours-scale restore and a 180-day minimum.
- Those minimums are billing commitments, not deletion locks. Deleting or transitioning early can leave a remaining-duration charge. This is why a short lab uses Standard even when an archive class has a lower advertised storage price. Storage-class comparison
Lifecycle, observation and bulk actions
- Lifecycle filters by prefix, tags or size, then transitions or expires matching objects. Treat current versions, noncurrent versions, delete markers and incomplete multipart uploads as separate cleanup concerns.
- In a versioned bucket, ordinary expiration can create a delete marker while old versions remain billable. Noncurrent-version expiration removes those older versions.
- Current lifecycle defaults do not transition objects smaller than 128 KB. Explicit size filters can change eligibility, but many tiny transitions may cost more than they save. Do not confuse the storage-class minimum billing duration with the object's age before a lifecycle transition is eligible. Lifecycle constraints
- Storage Class Analysis observes access patterns to inform Standard-to-IA choices; it does not perform the transition. Storage Lens aggregates usage/activity across buckets and accounts; advanced metrics can cost extra.
- Batch Operations uses a manifest and an execution role to apply supported actions across large existing object sets. Think “managed bulk job,” not “event notification for future writes.”
- Requester Pays makes authenticated requesters pay qualifying requests and downloads; the owner still pays storage. It is not anonymous public access and does not shift every possible charge.
Notifications and throughput
- S3 events can target SNS, SQS standard queues or Lambda; EventBridge adds filtering and more destinations. Direct S3 notifications cannot target SQS FIFO.
- Expect duplicate and potentially out-of-order notifications. Make processing idempotent; use separate subscriber queues when each consumer needs its own retry history. Avoid a function repeatedly triggering itself by writing back to its input prefix. Notification destinations
- Multipart upload sends large objects as independent parts, allowing parallelism and retrying only failed parts. Unfinished parts keep consuming storage until the upload is completed or aborted.
- Byte-range GET fetches selected bytes and can parallelize downloads. Transfer Acceleration improves long-distance client-to-S3 transfer paths through edge networking; evaluate its additional cost.
- S3 scales per partitioned prefix. Parallel requests and sensible prefix distribution improve throughput; KMS quotas can also constrain SSE-KMS workloads. Randomizing every object-name prefix is not a universal requirement. Performance guidance
See storage foundations for versioning, replication and durability versus availability, and S3 access controls for encryption and policy evaluation.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Predictable aging and expiration | Lifecycle |
| Unknown or changing object access | Evaluate Intelligent-Tiering |
| Rare data must be readable immediately | IA or Glacier Instant, depending on retention/access pattern |
| Millions of existing objects need one supported action | Batch Operations |
| Independent processing and audit consumers | Fan-out with separate queues |
| Distant clients upload large objects | Test acceleration and multipart upload |
| Investigate storage growth across accounts | Storage Lens |
Traps
- Cheap archival storage can be expensive for objects deleted tomorrow or retrieved frequently.
- Lifecycle is asynchronous; it is neither an exact timer nor a complete replacement for explicit teardown.
- A prefix is part of an object key, not a filesystem directory with independent throughput hardware.
- Delivering an event once and applying a business side effect once are different guarantees.
12 · S3 security
Memory hook: Encryption protects stored bytes, policies authorize callers, and browser rules do neither job for you.
Must remember
Match encryption to the key-control requirement
- SSE-S3: S3 manages encryption keys. New S3 objects receive server-side encryption by default; that baseline does not mean every bucket satisfies a requirement for a particular customer-managed key.
- SSE-KMS: KMS adds key-policy control and key-use auditing. A caller may need both S3 authorization and KMS authorization; an allowed object read can still fail when key access is missing.
- S3 Bucket Keys reduce eligible SSE-KMS calls and cost. They do not replace the KMS key policy or make unauthorized callers trusted.
- DSSE-KMS: two layers of server-side encryption for requirements explicitly calling for dual-layer protection. S3 Bucket Keys are not supported with DSSE-KMS. DSSE-KMS guidance
- SSE-C: S3 performs encryption but the customer supplies the key over HTTPS for relevant operations; S3 does not retain that key. Losing it can make the object unrecoverable.
- Current SSE-C caveat: since April 2026, new general-purpose buckets—and existing buckets in accounts without SSE-C objects—block new SSE-C writes by default. It requires deliberate enablement; these labs do not enable it. Client-side encryption is different: the client encrypts before uploading. SSE-C behavior
Separate defaults, enforcement and exposure
- Default encryption supplies a storage behavior. A bucket-policy deny can enforce a required encryption header/key; design conditions carefully because a missing header and an explicitly incorrect header are different requests.
- TLS enforcement uses a deny for insecure transport, commonly
aws:SecureTransport = false. Encryption at rest does not protect HTTP traffic. - Block Public Access is another independent safeguard. Identity policies, resource policies, explicit denies and applicable account controls still participate in authorization.
- CORS lets a browser expose responses across specified origins/methods. It is not an S3 permission and is not relevant to every non-browser client.
- Pre-signed URLs temporarily use the signer's permission. Access can end when the URL expires, the signing credentials expire, or authorization is revoked; the requested URL lifetime is not a guaranteed lifetime.
- Server access logs provide best-effort request records in a logging bucket. They are not a synchronous authorization gate or a complete substitute for selected CloudTrail data events. S3 security guidance
Retention and application-specific access
- Object Lock protects object versions, not merely a filename. Governance permits specifically authorized bypass; compliance cannot be shortened or bypassed during retention, including by root.
- A legal hold has no automatic expiry and remains until released by an authorized principal. It is independent of a version's timed retention. Either can prevent deletion.
- MFA Delete adds MFA requirements to selected versioning/permanent-deletion operations and requires root-controlled setup. It is different from Object Lock and is not configured here.
- Glacier Vault Lock fixes a retention policy for the separate legacy Glacier vault model. Do not confuse it with S3 Glacier storage classes or S3 Object Lock. Object Lock concepts
- Access Points provide separate endpoints and policies over shared bucket data; they do not create independent object copies or bypass the bucket's security controls.
- Object Lambda historically transformed S3 reads through Lambda. Since November 7, 2025 it is limited to existing users and selected partner solutions. Recognize the course pattern, but use supported application/edge transformation designs for new customers. Availability notice
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Control and audit use of a customer-managed key | SSE-KMS plus correct key/S3 policies |
| Explicit dual-layer encryption requirement | DSSE-KMS |
| Data must arrive at AWS already encrypted | Client-side encryption |
| Temporary download of one private object | Pre-signed URL |
| Authorized browser request blocked cross-origin | Inspect CORS |
| Non-bypassable retention of a version | Object Lock compliance |
| Separate application policies over one bucket | Access Points |
Traps
- Making CORS permissive cannot fix missing IAM or KMS permissions.
- Changing a bucket's default encryption does not retroactively re-encrypt all existing versions.
force_destroyis not stronger than an AWS retention lock.- A URL is a temporary bearer capability: anyone receiving it can use its allowed access while it remains valid.
13 · CloudFront and Global Accelerator
Memory hook: CloudFront caches HTTP content, Global Accelerator routes network connections, and replication creates durable regional copies.
Must remember
Match the origin and cache behavior to the application
- CloudFront accepts viewer HTTP(S) requests at edge locations. A distribution can use different origins and cache behaviors for static assets and dynamic API paths.
- S3 REST origin plus OAC: CloudFront signs origin requests, and a bucket policy permits the intended distribution. The bucket can remain private. For SSE-KMS objects, the key policy also needs appropriate access.
- An S3 website endpoint is a custom origin, cannot use OAC, and does not provide origin HTTPS. Do not choose it when the requirement is private S3 access through signed origin requests. S3 origin restrictions
- ALB/EC2 origins serve dynamic HTTP applications. CloudFront can pass requests through without caching user-specific responses; “CDN” does not mean static files only.
- Current VPC origins support eligible private-subnet ALBs, NLBs and EC2 instances, subject to service/Region constraints. Do not memorize the outdated rule that every application origin must be public. Public custom origins still need protection against bypassing the distribution. VPC origins
Keep caching separate from authorization
- The cache key decides which requests can reuse a response. Include only relevant headers, cookies and query parameters: unnecessary variation lowers hit rate; missing user-specific variation can expose private content.
- The origin request policy decides what reaches the backend. Forwarding a value and including it in the cache key are different decisions. Disable caching where that is the safest fit for personalized data.
- TTL bounds freshness. Versioned object names make immutable assets easy to update without replacing the contents behind a cached URL. Invalidations remove cached paths earlier and may add cost.
- Signed URLs suit individual restricted resources or clients that do not support cookies. Signed cookies suit access to multiple restricted files without rewriting each URL, such as a media presentation with many segments.
- OAC authorizes CloudFront to S3; signed URLs/cookies authorize viewers to CloudFront. They solve different boundaries and may be used together. Signed access choices
- Geo restriction controls viewer countries based on location signals. Price classes constrain the eligible edge footprint to trade price against reach. Neither determines the S3 bucket's storage Region or guarantees data residency.
- TLS applies on both viewer and origin connections. For a CloudFront custom hostname using ACM, the viewer certificate is obtained in us-east-1; an ALB's certificate is regional to that ALB.
Distinguish routing, resilience and edge code
- Global Accelerator provides static anycast IPs and routes TCP/UDP through AWS networking to supported healthy endpoints. Endpoint groups are regional; traffic dials and endpoint weights influence traffic distribution.
- It does not cache S3 objects or replace application authorization. Its control API uses us-west-2, although application endpoints can be elsewhere. Global Accelerator overview
- S3 CRR stores durable copies in another Region. CloudFront caches can expire or evict objects, so caching alone is not regional disaster recovery.
- CloudFront Functions is for lightweight viewer request/response logic, such as URL or header changes. Lambda@Edge supports richer processing and origin events, with different limits and replicated-resource lifecycle behavior.
- CloudFront changes/deletion need propagation. Lambda@Edge replicas can delay cleanup; avoid treating edge code as an ordinary instantly deleted regional function. See serverless services and S3 replication.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Millions of identical global HTTP downloads | CloudFront caching |
| Private S3 origin with controlled viewer access | OAC plus signed viewer access |
| Many restricted media segments | Signed cookies |
| Static global IPs for TCP/UDP applications | Global Accelerator |
| Durable regional recovery copy | S3 CRR |
| Personalized API behind an ALB | Appropriate cache policy or caching disabled |
| Simple viewer URL/header rewrite | CloudFront Functions |
Traps
- A CDN cache is not a backup, and a replica is not an authorization mechanism.
- Country filtering is not the same as choosing where original data is stored.
- Allowing all CloudFront traffic to a public origin does not necessarily restrict access to your specific distribution.
- A long TTL on a personalized response can be a security mistake, not just a freshness mistake.
14 · Storage extras
Memory hook: Choose the application's storage protocol first, then decide whether you need persistent access, hybrid access or data movement.
Must remember
Separate block, file and object needs
- EBS supplies block volumes with AZ placement constraints. An application normally creates a filesystem on the volume; it is not automatically a shared network filesystem.
- EFS supplies shared NFS access for Linux/POSIX workloads. Choose its availability and throughput configuration deliberately; “shared” does not mean every file-storage protocol is supported.
- S3 supplies an object API. Object keys are not ordinary disk blocks or a full POSIX filesystem. An application needing native file locking or SMB semantics usually needs another access layer.
- FSx is a family of managed specialized filesystems. Choose the correct family rather than treating every FSx option as interchangeable. AWS storage overview
Recognize the FSx families by their strongest clue
- FSx for Windows File Server: Windows SMB shares, NTFS permissions and Active Directory integration. This fits Windows home directories and applications that depend on Windows file-server behavior; EFS is not a drop-in replacement.
- FSx for Lustre: parallel high-throughput file access for HPC, analytics and machine learning. Link S3 datasets to a fast computation tier; understand scratch versus persistent deployment requirements rather than assuming every filesystem copy is your durable source.
- FSx for NetApp ONTAP: NetApp features and multiprotocol access, including NFS, SMB and iSCSI. Familiar snapshots, cloning, efficiency and migration capabilities matter when preserving an existing NetApp environment.
- FSx for OpenZFS: managed ZFS-oriented file workloads over NFS, with snapshots and cloning. The cue is ZFS compatibility and NFS behavior, not Windows SMB integration. Windows file-server features
- Across all four, check performance capacity, network placement, backup/restore and availability options separately. A fast scratch filesystem and a highly available production file store answer different questions.
Match a gateway to the on-premises interface
AWS Storage Gateway connects local applications to cloud storage while preserving a file, block or tape interface.
- S3 File Gateway exposes NFS/SMB file shares backed by S3 objects, with local caching. It is useful when an existing file-based workflow must use object storage without rewriting every client.
- Volume Gateway exposes iSCSI block volumes. Cached volumes keep primary data in AWS and cache frequently accessed data locally; stored volumes keep the full dataset locally and back it up to AWS.
- Tape Gateway presents a virtual tape library to existing backup software. It preserves the tape workflow while moving storage/archival responsibilities into AWS.
- A gateway provides an ongoing hybrid interface. It is not simply a one-time copying tool, and caching does not make network outages irrelevant.
- FSx File Gateway is no longer available to new customers. Do not confuse that restricted product with S3 File Gateway or FSx for Windows, which are separate services. Availability record
Distinguish transfer clients from migration jobs
- Transfer Family provides managed file-transfer capabilities such as SFTP, FTPS, FTP and AS2, with storage integration appropriate to the endpoint type. Use it when partners already speak a transfer protocol; it does not require them to adopt the S3 API.
- DataSync automates supported storage transfers, with task configuration, verification, scheduling and bandwidth controls. It is a strong fit for migrations and recurring synchronization rather than a general partner-facing SFTP login service. Agent requirements depend on the source/destination design. Transfer Family, DataSync
- Snow Family is the historical physical/offline transfer and edge-compute pattern: compare dataset size and available bandwidth before assuming a network transfer is feasible. Snowball Edge is closed to new customers; existing customers have separate availability and support timelines. This pack creates no physical jobs. Learn the architectural clue without assuming a new account can order devices. Product-specific availability notice
- Compare total cost: idle server/filesystem capacity, gateway infrastructure, transfer volume, requests and cross-AZ/Region traffic. “Managed” removes some operations work, not every charge.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Windows application requires SMB and AD | FSx for Windows |
| Parallel HPC processing of S3 datasets | FSx for Lustre |
| Preserve NetApp features and protocols | FSx for ONTAP |
| ZFS-oriented NFS migration | FSx for OpenZFS |
| Existing file shares need S3-backed hybrid access | S3 File Gateway |
| Existing backup software expects tape | Tape Gateway |
| Partners upload using SFTP | Transfer Family |
| Verified recurring bulk synchronization | DataSync |
Traps
- A protocol-compatible service can still have different availability, performance and feature limits.
- File Gateway, Volume Gateway and Tape Gateway preserve different client interfaces.
- A transfer endpoint moves or exposes data; it does not decide the authoritative storage and recovery strategy.
- EFS is not a universal Windows file share, and S3 is not a disk volume.
15 · SQS, SNS, Kinesis and Amazon MQ
Memory hook: Queue work for competing workers, fan out events to independent consumers, and retain streams when replay matters.
Must remember
SQS separates arrival rate from processing rate
- Standard queues provide at-least-once delivery and best-effort ordering. More workers can process different messages concurrently; producer and consumer availability no longer have to match exactly.
- Retention determines how long an unprocessed message can remain. Visibility timeout temporarily hides a received message while a worker processes it. Receiving does not delete it; delete only after successful processing.
- Set visibility around realistic processing/retry behavior. A worker can extend visibility for long work. If it crashes or visibility expires before deletion, another attempt can occur.
- Long polling waits for available messages, reducing empty receives and their cost. It does not extend retention or the worker's processing deadline.
- A DLQ receives messages after the configured receive-count failure threshold. Investigate and correct the cause before redriving; a DLQ is not successful completion.
- Queue resource policies authorize cross-service or cross-account senders, often with source ARN/account conditions. Encryption and permission to send are separate controls. Visibility behavior
FIFO ordering is scoped, and business effects still need protection
- FIFO queues preserve order within a message group. Different groups can progress independently; one global group limits parallelism.
- Deduplication IDs prevent duplicate sends within the five-minute deduplication interval. Content-based deduplication hashes the message body, not its attributes; repeated identical bodies need deliberate identity semantics.
- Deduplicated enqueueing is not a transaction with a payment provider or database. A worker may commit a side effect, crash before acknowledging, then receive the message again. Use an idempotency key and atomic/conditional business-state handling.
- Moving a message out to a DLQ can interrupt an application's intended sequence. If later messages must never overtake a failed operation, design the failure workflow accordingly. FIFO deduplication
SNS distributes copies; it does not create a worker backlog by itself
- SNS pushes messages to subscribers. Filter policies select messages by configured attributes or payload fields; they are not authorization policies.
- SNS plus SQS fan-out gives each subscriber its own copy, buffer, scaling and retry boundary. A single shared queue instead distributes work among competing consumers.
- SNS FIFO topics with compatible FIFO queue subscriptions support ordered fan-out; do not assume all subscriber types preserve the same ordering guarantees.
- S3 events can publish through SNS for independent consumers. Direct S3-to-SQS-FIFO notification is unsupported; EventBridge is a supported routing alternative. See object events. Messaging choices
Streams and broker compatibility answer different requirements
- Kinesis Data Streams retains records so independent consumers can read and replay them. A partition key maps records to a shard; ordering is scoped to the relevant shard/key, not the entire distributed workload.
- Provisioned mode means planning shard capacity; on-demand mode reduces that capacity-management work. Neither excuses a design that sends all traffic through one hot partition key.
- Consumers track processing position/checkpoints. Reading a record does not delete it for other applications. Retention and consumer lag determine whether replay remains possible. Kinesis Data Streams
- Amazon Data Firehose provides managed delivery into supported destinations, with buffering and optional transformation/format conversion. Choose it when delivery is the goal; choose a stream plus consumers when custom processing/replay is the goal. Firehose's buffering is not a general-purpose long-term replay contract. Firehose
- Amazon MQ manages supported ActiveMQ/RabbitMQ brokers. Existing JMS/AMQP-style applications that must preserve broker semantics may fit MQ better than rewriting around SQS/SNS.
- ASG worker scaling: approximate acceptable backlog per instance as acceptable queueing delay divided by average processing time. Scale on backlog per worker, not raw queue length alone; keep retries and downstream limits in the design.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Buffer bursts before independent jobs | SQS |
| Per-order sequencing with parallel unrelated orders | FIFO with a group per order/workflow |
| Every application must receive the event | SNS fan-out with separate queues |
| Replay telemetry through several consumers | Kinesis Data Streams |
| Managed streaming delivery into S3 | Firehose |
| Preserve an existing broker protocol | Amazon MQ |
| Scale workers while meeting queueing latency | Backlog per worker |
Traps
- Visibility, retention and delay are different clocks.
- FIFO does not make an external payment automatically exactly once.
- A filter decides delivery selection; a resource policy decides whether delivery is authorized.
- A partition key's ordering benefit can become a throughput bottleneck when too much work shares that key.
16 · Containers on AWS
Memory hook: Separate the container image, application task, compute capacity and permissions before choosing an orchestrator.
Must remember
Understand what each resource represents
- A Docker image packages application layers; a container is a running instance. Rebuilding an image and replacing running containers are separate operations.
- ECR stores images and versions/tags. Immutable tags prevent accidental tag replacement; digests identify specific content. Lifecycle rules remove old registry artifacts, but stopping tasks does not clean the registry.
- ECS task definitions describe containers, CPU/memory, roles, networking and logging. A task is an execution; a service maintains desired running tasks and replaces failures.
- ECS on EC2 leaves host capacity/patching and placement choices to your design. Fargate removes server provisioning while still requiring task sizing, subnets, security groups and suitable network access.
- An empty cluster or registered task definition is not running compute. Conversely, an idle-looking service with desired tasks can keep billing.
Learn the three ECS role boundaries
- Task role: permissions used by application code, such as reading S3 or DynamoDB. Give each workload only its required data/API scope.
- Execution role: actions the ECS/Fargate agent performs for the task, such as pulling a private ECR image and delivering configured logs.
- EC2 instance role: host/agent permissions for EC2-backed capacity. Do not use the host role as a substitute for per-task application permissions.
- A successful image pull does not prove the application can reach its database. IAM, routing, DNS, security groups and the data service's policies are separate checks. ECS IAM roles
Match scaling and storage to the workload
- ALB routes HTTP(S) to tasks and evaluates target health. Tasks using
awsvpcnetworking register as IP targets, not host instance targets. - Service auto scaling changes desired task count. Capacity scaling adds/removes EC2 hosts where needed. More desired tasks do not help if the cluster lacks capacity to place them.
- Fargate manages underlying capacity, but tasks still need valid CPU/memory combinations, quotas and reachable image/log endpoints. A private subnet may need NAT or appropriate service endpoints.
- EventBridge/Scheduler can start standalone tasks for a schedule or event. Choose this for intermittent work rather than an always-running service; target execution permissions and
iam:PassRolematter. - EFS supplies shared persistent files outside disposable containers. Mount targets, access points, permissions and NFS security-group paths matter. Container-local writable storage is not a shared persistent database.
- Separate application scaling from downstream limits: scaling containers can overload a fixed-capacity database. Service scaling
Recognize Kubernetes and hybrid variants
- EKS provides a managed Kubernetes control plane. It fits requirements for Kubernetes APIs, controllers and ecosystem compatibility, rather than merely “we have a container image.”
- Compute options include managed node groups, self-managed EC2, supported Fargate profiles and current managed options such as EKS Auto Mode. Responsibility and feature support differ.
- CSI storage drivers connect Kubernetes storage to AWS services. EBS is AZ-scoped block storage; EFS supports shared file access. Check the chosen compute/storage combination rather than assuming every volume works with every node type.
- ECS Anywhere runs registered external machines under the regional ECS control plane; you still manage those machines and connectivity.
- EKS Anywhere is customer-managed Kubernetes for supported on-premises/edge environments, including disconnected designs. EKS Distro is the Kubernetes component distribution, not a hosted control plane. EKS Hybrid Nodes instead connect customer-managed nodes to an AWS-managed regional control plane. EKS deployment choices, ECS external instances
- Remove controller-managed load balancers and persistent storage appropriately before removing Kubernetes controllers/cluster infrastructure, or external resources may be orphaned.
- App Runner historically simplified managed web-container deployment; App2Container analyzed/containerized existing applications. They are restricted for new customers as documented in the availability record; recognize their purposes without treating them as new sandbox defaults.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Containers with minimal host management | Fargate |
| Kubernetes compatibility | EKS |
| Specialized EC2 hosts or detailed capacity control | EC2-backed orchestration |
| Container code needs S3 access | Task role |
| ECS must pull a private image | Execution role |
| Occasional scheduled container job | Event-driven standalone task |
| Existing external machines under ECS control | ECS Anywhere |
| Customer-managed disconnected Kubernetes | EKS Anywhere |
Traps
- Task count, host count and Kubernetes control-plane availability are separate decisions.
- Giving a role permissions does not create a route or open a security-group path.
- Container-local state can disappear during replacement; scaling makes that weakness more visible.
- A distribution of Kubernetes software is not the same product as an AWS-managed Kubernetes cluster.
17 · Serverless services
Memory hook: Separate execution limits, data access patterns, API authorization and workflow state instead of treating serverless as unlimited capacity.
Must remember
Lambda execution, scaling and network access
- Lambda runs event-driven functions without provisioning ordinary application servers. Handlers should tolerate retries and avoid depending on one execution environment's local state.
- Concurrency is simultaneous execution, roughly arrival rate multiplied by execution duration. Faster functions can reduce concurrent capacity needed, but downstream services may remain the bottleneck.
- Reserved concurrency reserves capacity and caps that function; it does not pre-initialize environments. Provisioned concurrency keeps initialized environments ready and can incur idle cost. Account quotas, function scaling behavior and throttling still apply.
- The standard Lambda functions used in this pack have a 15-minute maximum invocation timeout. Current specialized execution modes have separate limits; do not turn that rule into a claim about every new Lambda feature. Memory also influences available CPU. Current Lambda quotas
- SnapStart restores eligible published functions from initialized snapshots. Check runtime/feature support and handle values that must remain unique or connections that can become stale after initialization.
- VPC attachment reaches private resources such as a private relational database. It does not give the function a public IP; a public subnet alone does not provide internet egress. Use appropriate NAT/routes or supported endpoints.
- Supported RDS/Aurora-to-Lambda integrations let database-side activity invoke a function with suitable engine, IAM and network configuration. That is different from Lambda opening a database connection.
- CloudFront Functions suits lightweight viewer logic; Lambda@Edge supports richer viewer/origin processing but introduces replicated-resource restrictions and slower cleanup. See global delivery.
DynamoDB access patterns, capacity and consistency
- A table has a partition key, or a composite partition key plus sort key. The partition key distributes data; the sort key orders related items within that partition-key value.
- Query uses a key condition to select a partition and optional sort-key range. Scan reads across the table/index. A filter is applied after reading; it does not transform an expensive scan into an efficient key lookup.
- On-demand suits variable demand without choosing provisioned throughput. Provisioned uses read/write capacity with optional auto scaling; predictable utilization can favor explicit sizing. Item size, read consistency and transactional operations affect capacity use.
- Capacity arithmetic: one RCU supports one strong read per second, or two eventual reads, for items up to 4 KB. One WCU supports one write per second for items up to 1 KB. Round larger items up to capacity-unit boundaries; transactions require additional capacity. Provisioned capacity
- LSI: same partition key, different sort key; define it with the table, share table capacity, and choose eventual or strong reads.
- GSI: a different access path using potentially different partition/sort keys; it can be added later and scales separately. GSI reads are eventually consistent, not strongly consistent. Index comparison
- Tables and LSIs support strong reads when requested. Strong reads cost more than equivalent eventual reads; neither means every future request is protected from subsequent writes.
- Global tables replicate across Regions. Multi-Region eventual consistency and multi-Region strong consistency are different modes with different support and tradeoffs; do not memorize “all global tables are eventually consistent.” Read consistency
- DAX caches suitable eventually consistent DynamoDB reads. It does not remove the need for good partition keys or make a strong-read requirement a cache hit.
- Streams captures item changes for event processing with per-item ordering; it is not permanent backup history. TTL deletes expired items asynchronously, so applications must handle records that have expired logically but remain physically present.
- PITR and on-demand backups restore into new tables. S3 export supports analysis of point-in-time data without consuming table read capacity; S3 import creates a new table. Replication and backup protect against different failures.
API Gateway and user identity
- REST APIs offer edge-optimized, regional and private endpoint types. HTTP APIs are a separate feature/cost choice; do not assume every REST feature exists on HTTP APIs.
- IAM authorization uses signed AWS requests. REST APIs support Cognito user-pool authorizers; HTTP APIs have native JWT authorizers. Lambda authorizers allow custom authorization logic. REST API keys and usage plans provide client identification/metering controls, not sufficient authentication by themselves. API comparison
- Cognito user pools provide user directories, sign-in and tokens. Identity pools exchange supported identities for temporary AWS credentials through roles, optionally including configured guest access. They solve different problems and may be combined.
- A valid token proves an identity claim; the application still must enforce ownership of the requested record. See serverless application design.
Workflow orchestration and reusable applications
- Step Functions models tasks, choices, parallelism, waits, retries and catches. Prefer it to implementing a long workflow through functions that poll or sleep.
- Standard supports durable, auditable workflows up to one year and bills state transitions. Its exactly-once workflow model does not make every external side effect immune to explicitly configured retries.
- Express supports short, high-volume workflows up to five minutes and bills execution/duration/memory. Asynchronous Express is at-least-once; synchronous Express is at-most-once. Select a model compatible with the work's idempotency and integration needs. Workflow types
- AWS Serverless Application Repository publishes and shares deployable serverless applications, commonly described with AWS SAM. It is a reusable application catalog, not a runtime, container registry or marketplace data feed. Review permissions and created resources before deployment. Repository purpose
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Reserve and limit a function's concurrent work | Reserved concurrency |
| Reduce eligible initialization latency | Evaluate provisioned concurrency or SnapStart |
| New query needs a different partition key | GSI, accepting its consistency model |
| Immediate consistent read of a base-table item | Strong table read |
| React to item changes | DynamoDB Streams |
| User sign-in tokens | Cognito user pool |
| Temporary direct AWS access for app users | Cognito identity pool |
| Long-running auditable coordination | Standard Step Functions |
| Reuse a packaged serverless application | Serverless Application Repository |
Traps
- On-demand capacity does not eliminate hot keys, service quotas or downstream limits.
- TTL is not an exact scheduler; a global replica is not a historical backup.
- An LSI cannot simply be added to an existing table; a GSI cannot provide strong reads.
- A VPC subnet labelled “public” does not give Lambda public-address connectivity.
- Reserving concurrency is different from paying to keep environments initialized.
18 · Serverless solution architectures
Memory hook: Authenticate the caller, authorize the operation, keep durable state outside functions, and cache only responses safe to share.
Must remember
A mobile records application needs more than a login screen
- A common request path is client → API Gateway → Lambda → DynamoDB. Cognito can provide sign-in and tokens; an authorizer validates identity at the API boundary.
- Authentication identifies the caller. Authorization decides whether that caller can perform this operation on this record. Derive the owner/tenant from validated identity rather than trusting an arbitrary user ID in the request body.
- Model DynamoDB partition/sort keys around the application's access patterns, such as “this user's records in date order.” Avoid reading everyone's data and filtering it in the client.
- User pools supply sign-in tokens. Identity pools supply temporary AWS credentials when clients need direct permitted AWS access. Do not ship permanent AWS access keys inside a mobile or browser application.
- Large uploads can go directly to S3 using a narrowly authorized pre-signed upload or temporary role credentials, while the API coordinates metadata and authorization. This avoids needlessly proxying all file bytes through application compute.
- API request throttling, Lambda concurrency limits and database capacity form one system: buffering or explicit backpressure can be more useful than increasing every quota. See serverless service choices. AWS serverless architecture patterns
A publishing website has static and dynamic paths
- Store immutable HTML/assets/media in S3 and serve them through CloudFront. Use a private S3 REST origin with OAC where the origin must be inaccessible directly.
- Keep authenticated comments, edits and personal data behind an authorized dynamic API. A public homepage and a private account page should not inherit the same caching assumptions.
- Choose cache keys around the representation being served. Missing relevant identity/authorization variation can expose private content; unnecessary cookies/headers can destroy useful cache reuse.
- Prefer long-lived caching for content-addressed or versioned assets. A short-lived manifest can identify the current release while older immutable assets remain cacheable.
- Separate viewer-to-CloudFront, CloudFront-to-origin, API-to-Lambda and Lambda-to-data permissions. A valid login does not automatically grant every downstream permission. API Gateway invocation permissions
Decouple microservices deliberately
- Let each service own a clear interface and data boundary. Sharing one writable database schema across every service can make deployment and failure isolation harder.
- Synchronous APIs fit operations that require an immediate answer, but the caller inherits downstream latency and failures. Use timeouts, limited retries and appropriate circuit-breaking/backpressure behavior.
- Queues/events fit deferred work. SQS buffers work; SNS/EventBridge can distribute events; independent queues let consumers recover at different rates. This introduces eventual completion rather than a synchronous transaction.
- Idempotency protects retries: identify the business operation and atomically recognize work already committed. A generated event ID alone is not enough if repeated requests get unrelated IDs.
- For multi-step work, Step Functions makes state, retry/catch paths and compensation visible. Compensating business actions are not equivalent to rolling back a distributed database transaction.
- Use metrics, structured logs and AWS X-Ray traces where supported to locate latency across service boundaries. A trace explains the request path; it does not authorize that request. See messaging patterns.
Deliver software and test real clients
- Software distribution: place versioned binaries in S3 and cache them at CloudFront edges. Calling Lambda for every identical download adds work without creating a CDN cache.
- Signed URLs restrict particular downloadable resources; signed cookies can authorize a collection without rewriting every asset URL. Keep OAC or other origin restrictions so clients cannot bypass the viewer control.
- CDN caches improve delivery, not durable recovery. Use replication and an appropriate origin recovery/failover design if regional resilience is required; measure RPO/RTO separately. Signed viewer access
- AWS Device Farm tests mobile/web applications on hosted real devices and browser environments. It addresses device/browser compatibility and user-interface behavior, not API hosting, application distribution or infrastructure provisioning.
- A backend unit test can pass while a particular mobile browser fails sign-in, CORS handling or rendering. Device testing and API/integration tests therefore complement each other; use controlled test accounts and datasets. Device Farm purpose
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Authenticated per-user records | Authorizer plus item-level authorization |
| Static global content with few origin reads | S3 and CloudFront |
| Direct large client uploads | Authorized direct S3 upload pattern |
| Slow background processing | Queue and asynchronous consumer |
| Multi-step retry/compensation workflow | Step Functions |
| Restricted collection of downloadable files | Signed cookies |
| Client compatibility across physical devices | Device Farm |
Traps
- A valid JWT does not grant permission to another user's record.
- Broadly caching personalized API responses can leak data even when the origin is private.
- “Serverless” does not remove quotas, idle provisioned features, request charges or service failure boundaries.
- A directly invocable Lambda function can still fail when API Gateway lacks its own invoke permission.
19 · Choosing an AWS database
Memory hook: Start with the required queries and consistency, then choose the data model, resilience and operating model.
Must remember
Ask the questions in a useful order
- Access pattern: known-key lookup, relational join, relationship traversal, full-text search, time-window aggregation or large-object retrieval?
- Correctness: what must be transactional, how current must a read be, and can users tolerate eventual consistency or stale cached data?
- Scale and latency: reads versus writes, predictable versus bursty demand, hot keys, record sizes, geographic access and tail latency.
- Resilience: AZ failure, Region failure and accidental data changes are different failure cases. Specify RPO/RTO and test recovery instead of selecting a service by a generic “highly available” label.
- Operations and migration: engine/API compatibility, managed versus self-managed responsibilities, required extensions, licensing and allowed application changes.
- Cost: include idle capacity, replicas, indexes, cache clusters, I/O, backups and transfer—not only the smallest instance price. “Serverless” and “NoSQL” are properties, not complete selection criteria. AWS database selection guide
Relational, keyed and cached data
- RDS fits a supported relational engine when SQL compatibility and limited application change matter. Aurora offers MySQL/PostgreSQL-compatible relational choices with its cluster-storage and endpoint architecture.
- Separate read scaling from failover availability. A read replica, a standby and a backup have different jobs; details depend on deployment type. Relational does not imply that every write can scale horizontally without redesign. See relational resilience.
- DynamoDB fits modeled key-value/document access patterns with managed scaling. Partition/sort keys and indexes must support the required queries; scanning a huge table for routine lookups is a design warning.
- On-demand capacity can fit unpredictable requests; provisioned capacity can fit predictable utilization. Strong reads, transactions, indexes and global replication have specific costs and behavior. “NoSQL” does not mean transactions are impossible. See DynamoDB keys and consistency.
- ElastiCache reduces repeated database work and latency. Valkey/Redis OSS and Memcached have different data structures, persistence/replication and operational capabilities; choose the engine/configuration rather than assuming every cache is interchangeable.
- Lazy loading/cache-aside: on a miss, read the authoritative store and populate cache. Write-through: update cache along with application writes, trading additional work for more immediately populated entries. Both need an invalidation/expiry strategy.
- A session store can move user state out of disposable web instances. Decide what happens when a cache entry expires or a node fails; some sessions can be reconstructed, while others require stronger persistence.
- S3 stores durable objects and data-lake files. Use it for large media or immutable data, often with keyed metadata elsewhere. It is not a drop-in transactional row database.
Purpose-built models
- DocumentDB provides MongoDB-compatible document APIs. It can fit JSON-document workloads, but compatibility is version/feature-specific: check operators, indexing behavior and drivers before assuming an unchanged migration.
- “Document” describes the data/query model; simply storing JSON does not require DocumentDB if a keyed DynamoDB item or an S3 object meets the actual access pattern. DocumentDB compatibility
- Neptune serves graph workloads built around connected entities and traversals, such as fraud relationships, identity links and recommendation paths. Prefer a graph model when relationship navigation is central, not merely because the dataset contains foreign keys. Neptune concepts
- Keyspaces provides managed Apache Cassandra-compatible wide-column access. The cue is an existing Cassandra/CQL ecosystem or a suitable distributed wide-column model, with service-specific compatibility constraints. Keyspaces concepts
- Time-series storage matches timestamped measurements, retention and time-window queries. Timestream for LiveAnalytics is closed to new customers; Timestream for InfluxDB is a distinct offering with a different operating model. Preserve the time-series concept without treating the names as interchangeable. LiveAnalytics availability
- QLDB historically supplied an immutable, cryptographically verifiable journal under a central owner. Support ended July 31, 2025. Recognize the historical integrity/ledger pattern, but choose a supported current design for new systems. See the availability record.
Transactional stores are not every analytical tool
An application's primary database handles its operational reads/writes. Large historical aggregation may belong in S3/Athena or Redshift; relevance search may need OpenSearch. Copying suitable data into a specialized read/analytics system can protect the transactional workload, but introduces freshness and pipeline responsibilities. See analytics choices.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Existing relational application with joins | Compatible RDS/Aurora engine |
| Known-key low-latency application records | DynamoDB |
| Repeated hot reads or external session state | ElastiCache with explicit miss/recovery behavior |
| Large media objects | S3, often with database metadata |
| MongoDB-compatible document queries | Evaluate DocumentDB compatibility |
| Multi-hop fraud/relationship traversal | Neptune |
| Existing Cassandra/CQL workload | Keyspaces |
| Timestamped measurements and time windows | Supported time-series design |
| Large analytical aggregation | See Athena/Redshift choices |
Traps
- A replica can copy an accidental deletion; it is not a historical recovery point.
- A cache hit can be fast and stale at the same time.
- API compatibility is not a promise that every engine feature or performance characteristic matches.
- Choosing one database for every workload can increase complexity rather than reduce it.
20 · Data and analytics
Memory hook: Separate ingestion, storage, metadata, permissions, computation and visualization so each requirement has a clear owner.
Must remember
Catalog and govern the lake before choosing the dashboard
- S3 holds data-lake objects. Glue Data Catalog holds metadata describing datasets, locations, schemas and partitions; a catalog entry does not contain or transform all of the underlying data.
- Glue crawlers discover supported schemas/partitions; Glue ETL jobs clean, transform and move data. Schema discovery and transformation are separate compute activities with separate costs.
- A known small schema can be declared directly instead of running a crawler. A newly declared column does not populate missing values in existing JSON objects.
- Lake Formation centrally governs lake access, including supported fine-grained permissions over cataloged data. IAM, storage access and key permissions remain relevant; enabling governance does not mean every existing access path is automatically secured. Glue concepts, Lake Formation
Athena and Redshift are different query choices
- Athena is serverless SQL for supported data sources, especially S3. For occasional queries, it avoids keeping a warehouse cluster running purely to wait for work.
- Partitioning allows queries with appropriate predicates to skip irrelevant data. Columnar formats such as Parquet/ORC, compression and selecting only needed columns reduce scanned work.
- Many tiny files add overhead; poor partition choices can also hurt performance. A
LIMITor a post-read filter is not a universal guarantee of low scan cost. - Workgroups isolate settings and controls such as query output location and scan limits. Keep result objects outside the table's source-data prefix so later scans do not mix source JSON with generated result formats.
- Federated queries access supported non-S3 sources through connectors. Some connectors introduce Lambda execution, spill storage and source-system load; “serverless SQL” does not make all dependencies free. Athena data optimization
- Redshift is an analytical warehouse suited to substantial joins, aggregations and recurring BI workloads. Columnar/parallel execution and table distribution/sort design matter; simply adding nodes is not every query optimization.
- Redshift Spectrum queries external S3 data without first loading it into Redshift tables. This can combine warehouse data with a larger lake rather than copying every cold dataset into the warehouse.
- Snapshots and cross-region copies provide recovery options. Storage, encryption/key access and restore time still matter; an available snapshot does not prove the required recovery time.
- Provisioned versus serverless deployment changes operations and billing, not the need to size, govern and monitor the analytical workload. Spectrum
Processing, search and streaming
- EMR runs frameworks such as Spark and Hadoop with supported cluster/serverless execution choices. Choose it when the workload needs that processing ecosystem or detailed framework control; idle cluster capacity can remain costly.
- OpenSearch indexes data for relevance search, aggregations and log analytics. An index is not automatically the authoritative transactional record; ingestion, retention and recovery need design.
- MSK supplies managed Kafka-compatible infrastructure. It fits existing Kafka clients, tooling and stream semantics; it does not by itself execute all application analytics.
- Managed Service for Apache Flink runs stateful streaming computations with windows, event-time handling and checkpoints. It can consume streams and maintain continuous results rather than rerunning isolated batch queries.
- Kinesis Data Streams provides a retained stream for independent consumers; Firehose provides managed buffered delivery. Choose the ingestion/replay and delivery functions separately. See messaging and stream choices.
- The old Kinesis Data Analytics for SQL service is retired. That is not the same product as current Managed Service for Apache Flink. Flink concepts, availability record
Current BI naming and external datasets
- The current exam service list names Amazon Quick. AWS describes Amazon Quick Sight as the continuing BI/visualization capability within Quick; older material calls it Amazon QuickSight. Preserve that name mapping rather than assuming the dashboard capability disappeared.
- For the familiar exam pattern, choose Quick Sight/QuickSight for interactive dashboards, business-user visualization and embedded analytics. SPICE stores imported analytical datasets in memory; refresh behavior determines when new source data appears. Direct query and imported data have different performance/freshness tradeoffs.
- Quick also includes broader research, automation, connected knowledge and application features. Its existence does not turn a dashboard into the database or ETL engine. No subscription is created for these labs. Current Quick terminology
- AWS Data Exchange manages access/entitlements to externally shared datasets and marketplace data products. It fits acquiring or sharing third-party data, not converting file formats or operating a query engine.
- Data grants/subscriptions, dataset revisions and allowed use must be understood before incorporating external data into analytics. Receiving a dataset does not automatically catalog, cleanse or visualize it. Data Exchange
Trace one complete pipeline
A plausible design is Kinesis/MSK → Flink or Firehose → S3 → Glue/Lake Formation → Athena/Redshift → Quick Sight. These are choices, not mandatory boxes: a simple batch dataset may need only S3, a catalog, Athena and a dashboard. Specify latency, duplicate handling, failure recovery, schema evolution and access controls at every handoff.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Occasional SQL over S3 files | Athena |
| Repeated warehouse joins and analytical reporting | Evaluate Redshift |
| Query cold S3 data from Redshift | Spectrum |
| Discover schemas or transform data | Glue crawler or ETL, respectively |
| Govern lake permissions centrally | Lake Formation |
| Preserve Kafka ecosystem | MSK |
| Continuous stateful window aggregation | Managed Service for Apache Flink |
| Search relevance over indexed documents | OpenSearch |
| Dashboards and embedded BI | Amazon Quick Sight, formerly QuickSight |
| Acquire/manage external dataset access | Data Exchange |
Traps
- Metadata, data files, query execution and visualization are separate layers.
- “Real time” requires a latency definition; buffered delivery and stateful event processing are not identical.
- A cache/imported BI dataset can be fast but stale until refreshed.
- Adding a schema field is not an ETL transformation, and an S3 bucket policy alone is not a complete lake-governance design.
- Serverless query/processing choices can still incur scans, capacity minimums, storage, request and transfer charges.
21 · Machine Learning
Memory hook: Identify the input and the required output before choosing an AI service.
Must remember
Words, voices and conversations
- Transcribe converts speech to text; Polly converts text to speech. Meeting subtitles point to Transcribe; reading a message aloud points to Polly. Translate changes the language of text, so translated subtitles can require Transcribe followed by Translate.
- Comprehend interprets text: sentiment, entities, key phrases and language insights. Comprehend Medical specializes in supported clinical entities and relationships; extracting a medication name does not make it a diagnostic system.
- Lex manages conversational intents, slots and dialogue. Connect supplies contact-center capabilities and telephony, now documented as Connect Customer. A support flow can use Connect for the call, Lex to understand the request and backend services to fulfill it. A chatbot and a complete contact center are different requirements.
- Distinguish language translation from sentiment classification. An application may need both, but one API response does not automatically supply the other. Confirm language and feature support for the particular operation rather than assuming every language works everywhere.
Images, documents and recommendations
- Rekognition analyzes images and video: visual labels, supported face analysis and other image-understanding features. Textract extracts document text and structure, including supported forms and tables. When the required result is an invoice's fields and line items, structured document extraction is the stronger clue than the generic word “image.”
- Personalize creates individualized recommendations using interaction and item information. It is not an enterprise document search engine. Kendra searches indexed enterprise content; the index and connector lifecycle add cost and administration. Both are retained from the requested curriculum even though neither is explicitly named in the current in-scope service list.
- SageMaker AI provides custom model development, training and deployment. Choose it when the workload needs a model or workflow that a suitable pretrained API does not already provide. Owning training data, evaluation and model serving adds decisions about quality, capacity and operations.
Architecture decisions around the API
- A synchronous request fits a small interactive operation; larger supported jobs may need asynchronous completion, an output store and retries. Distinguish the job's lifecycle from the data it writes. Deleting a job does not necessarily remove stored output.
- Keep application credentials in a service role. Limit actions and data access, encrypt sensitive input/output, and select a Region that meets data-handling requirements. Medical text and face data need particularly deliberate handling; use harmless sample inputs in this pack.
- Decouple bursty work with a queue and handle throttling with bounded retries. Evaluate accuracy against the actual business task; a confidence score is not a guarantee of correctness. Include a human review step when uncertain predictions cannot safely drive an automatic action.
- Pay attention to persistent capacity: an occasional API call, a running inference endpoint, a notebook and an enterprise-search index have different idle-cost behavior. None is justified merely because it is available in a console.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Audio recording needs searchable text | Transcribe |
| Written directions must be spoken | Polly |
| Reviews need sentiment scores | Comprehend |
| Invoice needs named fields and table cells | Textract |
| An image needs visual labels | Rekognition |
| Voice support needs intents plus call handling | Lex plus Connect |
| A bespoke model must be trained and hosted | SageMaker AI |
| Shopping suggestions depend on each user's history | Personalize concept |
Traps
- The service name is not enough: match supported input, output, language, latency and quality requirements.
- Reading text inside an image is not automatically equivalent to extracting document structure.
- “Managed” does not mean “no persistent resources” or “no idle charge.” Delete endpoints, jobs and outputs according to their separate lifecycles.
22 · Monitoring & Audit
Memory hook: Metrics show symptoms, logs explain events, traces follow requests, and audit records identify changes.
Must remember
Separate the evidence questions
- CloudWatch: how is the workload behaving? Use metrics, logs, dashboards and alarms for operational evidence. CloudTrail: which identity called which API, against what resource and when? AWS Config: what resource configuration existed, and did an evaluated rule consider it compliant?
- These sources complement each other. A slow API can require a latency alarm, application logs and a trace; identifying an administrator's change requires audit evidence. Check that the needed events, resources and retention were actually configured before promising historical answers.
- AWS X-Ray follows instrumented requests across services and downstream calls. Its trace map helps find latency, errors and bottlenecks. Traces are not a replacement for every application log or API audit event; instrumentation and sampling affect visibility.
Metrics and logs
- A metric is a numeric time series identified by namespace, name and dimensions. Choose meaningful statistics and evaluation periods: average latency can conceal slow tail requests, while a total error count without request volume may mislead.
- Alarms evaluate metric conditions; actions notify or invoke supported responses. Composite alarms combine alarm states to reduce noisy paging. Treat missing data intentionally rather than assuming missing means healthy. Metric streams continuously deliver selected metric updates to downstream consumers.
- Logs Insights queries log events. Metric filters count matching events into metrics. Subscriptions forward matching logs to supported destinations; export writes log data to S3 for a different processing workflow. These are distinct mechanisms.
- The CloudWatch agent collects additional guest-OS and application signals. Standard EC2 metrics do not automatically reveal every filesystem or memory measurement.
- Container Insights and Lambda Insights add workload-specific visibility; Contributor Insights identifies prominent contributors; Application Insights helps correlate application problems. Their scope and collection costs need deliberate configuration.
Events, audit and compliance
- EventBridge rules match events on buses and deliver them to targets. Scheduler invokes targets on time-based schedules. An archive retains selected events for replay; a replay can repeat a business action, so consumers still need idempotency.
- A successfully accepted event can fail to match a rule or fail later delivery. Check source/detail pattern, bus, target configuration, resource permissions and retry/dead-letter behavior separately.
- CloudTrail management events describe control-plane activity; selected data events provide supported resource-level activity. CloudTrail Insights detects unusual supported API activity. A rule reacting to an API call via EventBridge needs the appropriate event path and coverage.
- Config recording tracks selected resource configurations; rules evaluate compliance. Notifications and remediation are separately configured. A remediation role can mutate resources, so evaluation should not be confused with automatic repair.
Additional published-scope tools
- Amazon Managed Service for Prometheus stores and queries compatible operational metrics, especially for container workloads using PromQL. Amazon Managed Grafana visualizes metrics, logs and traces from multiple sources. The dashboard layer is different from the metric storage/query layer.
- AWS Health Dashboard reports AWS service events and account-relevant impacts. Combine it with workload telemetry: a healthy AWS status does not prove your application or configuration is healthy.
- Keep logs and traces useful: redact sensitive data, set retention, scope collection and correlate request identifiers. Broad logging can create both sensitive-data exposure and substantial ingestion charges.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Alert on errors or latency | CloudWatch metric alarm with an action |
| Determine who deleted a database | CloudTrail audit events |
| Review a resource's historical configuration compliance | Config history and rule evaluations |
| Find the slow downstream call in a distributed request | X-Ray tracing |
| Query container metrics using PromQL | Managed Service for Prometheus |
| Visualize several telemetry sources together | Managed Grafana |
| Trigger work for matching application events | EventBridge rule and target |
| Investigate a relevant AWS service disruption | AWS Health Dashboard |
Traps
- An alarm with no action is not an email subscription; rule creation alone does not establish every permission needed for delivery.
- A trace sample or a log metric is not a complete security audit record.
- A Config rule can report a problem without correcting it. Replaying an event can repeat side effects rather than merely replaying a picture of history.
23 · IAM Advanced
Memory hook: First identify the caller, then find its grants, its permission ceilings and every applicable explicit deny.
Must remember
Follow the permission decision
- Identity policies grant actions to users, groups and roles. Permissions boundaries limit the permissions that identity policies can grant. Service control policies (SCPs) set organization-level permission ceilings for affected member accounts. Neither a boundary nor an SCP grants access merely by allowing an action.
- SCP scope: member-account principals, including the member account's root user, are affected; management-account principals and service-linked roles are not. An account's SCP also does not constrain an external account's principal merely because it accesses that account's resource. Resource-access policies and controls must address that separate boundary.
- Explicit deny wins. For ordinary identity-policy access, a grant must survive the applicable boundary, session policy and organization restrictions. Do not turn that mnemonic into a universal set-intersection algorithm: resource-policy grants to users, role principals and role-session principals have different evaluation details.
- Role trust answers who may assume the role; the role's permission policies answer what the resulting session can do. Cross-account assumption normally needs suitable caller authorization and target trust. STS provides temporary credentials rather than a new long-lived access key.
- A resource-based policy can authorize a supported resource directly; assuming a role changes the effective identity used for subsequent operations. Choose according to the service and access pattern. Cross-account access also requires the applicable permissions on both sides.
- IAM simulation is useful evidence for supported policy evaluation, but it does not fully test every trust, organization, resource-policy or runtime context. A successful simulation of an action is not proof that MFA-conditioned role assumption will succeed.
Conditions carry request context
- aws:SourceIp constrains supported source-IP contexts. Calls through services or private endpoints can have different context; a public source-IP restriction is not a universal VPC restriction.
- aws:RequestedRegion tests the Region of the requested endpoint. Account for global services and cross-Region side effects; it is not automatically a guarantee that all resulting data stays in that Region.
- aws:PrincipalOrgID constrains supported access by organization membership. It can reduce the need to enumerate account IDs in suitable resource policies, but does not grant every principal the intended action.
- MFA conditions depend on how credentials and sessions were obtained. Do not assume a federated sign-in with MFA always supplies aws:MultiFactorAuthPresent in subsequent requests. A missing condition key matters to the selected policy operator.
- Least privilege includes actions, resources and conditions. Keep the operator identity separate from the role being tested so an exercise cannot silently weaken your own account access.
Multi-account and workforce choices
- Organizations groups accounts; OUs organize accounts for governance; SCPs constrain permissions. Separate production, development and shared services where isolation and governance require it. Organization policies do not replace each account's grants or data policies.
- IAM Identity Center provides workforce access through permission sets and account assignments, commonly backed by an external identity provider and temporary credentials. Consumer signup belongs to a different pattern, such as Cognito; see application identity.
- Directory Service: Managed Microsoft AD provides managed AD capabilities and supported trusts; AD Connector proxies authentication to an existing directory; Simple AD supplies a smaller compatible feature set where available. “Already have AD” is a clue to compare integration requirements rather than create another independent user database.
- AWS Resource Access Manager (RAM) shares supported resources across accounts, organizations/OUs or supported principals. A centrally owned subnet or Transit Gateway can be shared instead of duplicated. Resource sharing does not give participants universal ownership or bypass their IAM restrictions; supported resource types and sharing rules differ.
- Control Tower establishes governed multi-account landing zones with controls and integrated services. It is broader than one role or SCP. This pack treats Organizations and Control Tower as design concepts and leaves account-wide governance unchanged.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Employees need consistent temporary access across accounts | Identity Center permission sets and assignments |
| Developers create roles but must stay inside a maximum scope | Permissions boundaries plus controlled delegation |
| Restrict allowed services across an OU | SCPs, alongside actual identity grants |
| Share an existing supported network resource across accounts | RAM |
| Define who can assume an application role | Role trust policy |
| Existing directory must authenticate supported AWS workloads | Compare AD Connector and Managed Microsoft AD capabilities |
| Build a governed multi-account landing zone | Control Tower concept |
Traps
- Adding another Allow cannot override an explicit Deny or a limiting permissions boundary.
- A role ARN in a trust document is not the same thing as a resource permission granting that role every operation.
- Organization membership, network origin and MFA are different request properties; one condition does not establish the other two.
24 · Security & Encryption
Memory hook: Protect the connection, the stored data, the permission to use it and the evidence of misuse separately.
Must remember
Encryption and key control
- TLS protects transit; encryption at rest protects stored data. Neither prevents an already-authorized compromised application from reading plaintext. Authentication, authorization, secret handling and monitoring remain necessary.
- KMS: AWS-owned keys are managed within services; AWS-managed keys are visible in your account but have service-controlled administration; customer-managed keys provide your own policy/lifecycle control. Symmetric encryption, asymmetric operations and HMAC keys solve different cryptographic tasks.
- A KMS key policy is central to authorization. IAM permissions alone are not a universal substitute for a suitable key policy. Supported rotation keeps older material available for decrypting existing ciphertext; rotating a key does not automatically re-encrypt every stored object.
- Multi-Region KMS keys share related key material, enabling supported regional cryptographic use, but policies, grants, aliases and lifecycle remain regional decisions. Creating a replica does not copy every administrative setting or automatically replicate application data.
- Encrypted snapshot/AMI sharing needs resource permissions and appropriate customer-key access for the recipient. S3 replication of SSE-KMS objects needs explicit replication configuration, source decryption and destination encryption permissions with the correct destination key. A generic S3 copy policy is insufficient.
- CloudHSM provides dedicated hardware security modules and more direct cryptographic control, with greater administration/capacity responsibility. Choose it for a requirement that specifically needs that control or interface; ordinary managed encryption requirements often fit KMS better.
Configuration, secrets and certificates
- Parameter Store provides hierarchical configuration, Standard/Advanced tiers and KMS-backed SecureString. Secrets Manager supplies secret versions, supported rotation workflows and optional regional replication. Rotation requires the relevant integration and permissions; merely storing a secret does not rotate a database password.
- Applications should retrieve secrets using a role, with caching and refresh behavior appropriate to rotation. Terraform's sensitive flag controls some display behavior; it does not encrypt local state or prevent an authorized reader from recovering supplied values.
- ACM manages certificates. An ALB uses a certificate in its Region; CloudFront's ACM certificate must be in us-east-1. Validate domain ownership and consider the whole client-to-edge-to-origin TLS path rather than securing only one connection.
- AWS Private CA is an adjacent distinction: it issues certificates for a private trust hierarchy, such as internal services. Private certificates are not automatically trusted by public browsers, and a private CA introduces charges. It is not separately named in the current in-scope list, unlike ACM.
Filtering, detection and investigation
- WAF filters supported HTTP requests with web ACLs, IP sets and rules, including rate-based rules. Shield Standard supplies baseline DDoS protection; Shield Advanced adds paid capabilities. Firewall Manager centrally manages supported security policies across an organization. None replaces least privilege or secure application logic.
- DDoS resilience combines edge absorption, caching, rate controls, suitable scaling and protected origins. Keep expensive origin work from being the first line of defense. Network Firewall handles network inspection; WAF targets supported web request paths.
- GuardDuty detects suspicious activity. Inspector finds vulnerabilities in supported workloads. Macie discovers sensitive data in S3. Select based on the finding needed, not the generic word “security.”
- Security Hub, including its security-posture capabilities, consolidates findings and evaluates supported security controls. Detective helps investigate relationships and activity surrounding suspicious behavior. Aggregating a finding, investigating it and automatically remediating it are different steps.
- Artifact provides AWS compliance reports and agreements. It does not certify your application's configuration. Audit Manager can collect and organize evidence for assessments; it is useful adjacent context rather than an explicitly named service in the current list, and it does not replace the auditor's judgment.
- Modern Inspector remains relevant; Inspector Classic is retired. Consult service status for generation-specific dates. Never infer that a current service is unavailable solely because an older namesake ended support.
Operational boundaries
- Shared responsibility changes with the service: AWS operates underlying infrastructure, while you still control data classification, identities and workload configuration. Managing EC2 also includes guest-OS responsibilities that a fully managed service takes off your hands.
- Choose retention deliberately. Customer-key deletion has a waiting period; secret recovery settings and replicas affect deletion; immutable compliance retention can intentionally prevent removal. Such retention is valuable when required by a real workload and incompatible with this disposable lab's default.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Managed encryption with controlled key permissions | Customer-managed KMS key and appropriate policies |
| Dedicated HSM control or required cryptographic integration | CloudHSM |
| Automatically rotate supported database credentials | Secrets Manager with configured rotation |
| Sensitive information found in S3 objects | Macie |
| Vulnerable supported packages or images | Modern Inspector |
| Suspicious account/workload activity | GuardDuty |
| Consolidated findings and posture checks | Security Hub |
| Investigate connected security events and entities | Detective |
| Obtain AWS's compliance documentation | Artifact |
| Block abusive HTTP requests at CloudFront | WAF rules/IP sets/rate controls |
Traps
- Encryption is not authorization, and a resource share without key access can remain unusable.
- A managed certificate is not a domain registration; a private CA certificate is not automatically public trust.
- A detection service is not automatically a remediation engine. Enabling broad scans or organization controls can change account behavior and spending.
25 · Networking: VPC
Memory hook: A working connection needs the right address, a forward route, permission and a return path.
Must remember
Addresses and routes come first
- A VPC is a regional network boundary; a subnet occupies one AZ. Plan nonoverlapping CIDRs before connecting VPCs and on-premises networks. AWS reserves five IPv4 addresses in an ordinary subnet, so a /24 supplies 251 usable IPv4 addresses. Default-VPC conveniences should not be assumed in a custom VPC.
- A route table selects the most specific matching destination route. A subnet is called public when it has an internet-gateway route, but an IPv4 instance also needs usable public addressing and suitable security rules. A public IP without the route, or the route without the public IP, is insufficient.
- An internet gateway (IGW) supports the VPC's internet path. A bastion is a deliberate administrative hop; Session Manager can avoid inbound SSH by using the managed agent's outbound service connectivity and IAM permissions.
- Private addressing does not itself guarantee isolation from every network. Check all routes: peering, transit, VPN and service endpoints may create intentional private connectivity.
Stateful and stateless filters
- Security groups use stateful allow rules on interfaces/resources. Return traffic for an allowed connection is tracked; there is no explicit SG deny rule. Referencing another SG is useful for tier-to-tier access without maintaining individual IP lists.
- NACLs apply ordered stateless allow/deny rules at subnet boundaries. Both directions need appropriate rules; HTTPS responses often need outbound ephemeral ports. Lower-numbered matching rules determine the decision, so an earlier deny can defeat a later allow.
- A permitted filter cannot compensate for a missing route, and a correct route cannot override a denying filter. Trace the actual source/destination seen at each network hop.
Egress and service endpoints
- NAT instances require operating-system management, routing, security rules and appropriate source/destination-check changes. NAT gateways reduce appliance management; time, processing and associated address charges still matter. NAT permits outbound-initiated connectivity, not unsolicited inbound sessions.
- The traditional zonal public NAT gateway sits in a public subnet; same-AZ routing with one per required AZ avoids a single-AZ egress dependency. Current AWS also offers regional NAT gateways, which can automatically expand across AZs and do not require a hosting public subnet. Know which mode a question describes; regional mode currently does not provide private NAT. A single regional resource is not a promise of one-AZ pricing.
- Gateway endpoints for S3 and DynamoDB add route-table targets without an endpoint-hour fee. They are not general transit access for clients in peered VPCs or on-premises networks.
- Interface endpoints and AWS PrivateLink provide private access to supported services through endpoint networking and DNS, typically using private ENIs and SGs for interface endpoints. They can expose a particular service instead of granting full VPC-to-VPC routing. Hourly/per-AZ and data charges require comparison against the actual traffic pattern.
- Endpoint policies, where supported, limit use through that endpoint. They do not override missing IAM permissions or a denying bucket/resource policy. Network reachability and API authorization are separate checks.
Connect networks and resolve names
- VPC peering connects compatible nonoverlapping networks with explicit routes; it is not transitive. A–B and B–C do not establish an A–C path through B. Transit Gateway supplies a routed hub for many VPCs and on-premises connections; attachment and route-table configuration still control permitted paths.
- Site-to-Site VPN connects networks using encrypted tunnels over IP connectivity. VPN CloudHub supports compatible hub-and-spoke VPN site communication. Client VPN supplies remote-user access, with authentication, authorization and routes; it is not the same workload as linking two corporate networks.
- Direct Connect provides dedicated connectivity and more predictable network characteristics, with physical provisioning considerations. It is not encrypted by default. Use appropriate application TLS, supported MACsec or VPN designs when encryption is required. Direct Connect Gateway connects eligible virtual-interface designs to multiple VPCs/Regions; it is not an automatic transitive VPC router.
- Route 53 Resolver inbound endpoints let external networks query supported AWS DNS namespaces. Outbound endpoints and forwarding rules send matching VPC queries toward external DNS. Direction follows the query, and DNS resolution still needs underlying network reachability. See DNS notes.
IPv6, observation and cost
- Amazon-provided public IPv6 addresses are globally routable; routing and filters control reachability. An egress-only IGW permits outbound-initiated internet flows for public IPv6 addresses. Reaching IPv4-only destinations from IPv6 is a separate translation requirement.
- AWS also supports private IPv6 through IPAM, including ULA and private GUA ranges. IGWs and egress-only IGWs drop these private ranges; internet access requires a suitable intermediary with public addressing. “Outbound-only” and “private IPv6 address” are different properties.
- Flow logs summarize supported IP flows, including accepted/rejected traffic, rather than packet payloads. Delivering to S3 and querying with Athena supports analysis. Traffic mirroring copies supported packet traffic for inspection; Network Firewall provides managed network filtering/inspection. WAF specializes in supported HTTP request paths.
- Add the whole path's cost: public IPv4, NAT processing, cross-AZ/Region transfer, endpoints, load balancers and inspection. Service quotas, subnet address capacity and standby-region limits can prevent scale-out even when an architecture diagram looks sound.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Private VPC workloads only need same-Region S3 | Gateway endpoint and suitable policies |
| Expose one supported private service to another account | PrivateLink rather than broad routed connectivity |
| Many VPCs need transitive routing | Transit Gateway |
| Remote employees need authenticated private access | Client VPN |
| On-premises DNS must query private AWS names | Resolver inbound endpoint plus connectivity |
| VPC clients need corporate DNS zones | Resolver outbound endpoint and rules |
| Outbound-only internet access using public IPv6 addresses | Egress-only IGW |
| Inspect actual packet content | Suitable mirroring/inspection design |
Traps
- A subnet name, SG rule or public IP alone does not establish a complete path.
- A gateway endpoint is not a replacement for all interface endpoints or for hybrid connectivity.
- Older “all NAT gateways are zonal” shorthand is incomplete. Preserve the availability mode and routing assumptions in the question.
26 · Disaster Recovery & Migrations
Memory hook: RPO limits how much data may be lost; RTO limits how long useful service may be unavailable.
Must remember
Objectives before recovery patterns
- Recovery point objective (RPO) measures acceptable data loss as a time interval. Recovery time objective (RTO) measures acceptable recovery duration. Fast restoration cannot recreate transactions that never reached the available recovery data.
- Measure detection, decision, provisioning, data recovery, dependencies, traffic switching and validation. A database restore benchmark alone does not prove users can work within the RTO.
- Backup and restore: retain recovery data, then rebuild as needed. Pilot light: keep the minimal core running and add missing serving capacity. Warm standby: maintain a functional reduced-capacity workload and scale it. Multi-site active: serve from multiple sites while managing routing, consistency and failure behavior.
- Recovery patterns trade standing cost and complexity against recovery work; their names do not guarantee an RPO/RTO. Select using tested behavior and actual requirements.
- Availability across AZs and regional disaster recovery solve different failure scopes. Multi-AZ replication does not automatically protect against a Region-wide outage, accidental deletion or corrupted application data.
Backups need more than a schedule
- AWS Backup plans define rules; resource selections determine what is protected; vaults contain recovery points. A plan with no selections schedules no useful protection for the intended resource.
- Cross-Region or cross-account copies require compatible resource support, IAM and encryption-key arrangements. Verify the copy, its retention and the destination's ability to restore. A copied recovery point that cannot be decrypted is not a working recovery plan.
- Vault Lock provides retention controls; compliance-mode protection can intentionally prevent early deletion. Governance controls and compliance immutability have different bypass properties. This pack does not create retention that obstructs immediate cleanup.
- Keep historical recovery points where required: replication can quickly copy an unwanted deletion or bad write. Backup retention, replication and application validation protect against different failure modes.
- Validate service quotas, instance availability, subnet addresses, dependencies and key access in the standby Region before a disaster. A successful Terraform plan does not reserve all future capacity or raise every quota.
Migrate with a deliberate cutover
- DMS supports data movement with full load and change data capture (CDC) for supported endpoints. SCT helps convert schema/code for heterogeneous migrations; data replication does not automatically convert every stored procedure or engine feature.
- A low-downtime database migration commonly uses initial load, ongoing CDC, validation, a controlled write cutover and a rollback decision. Monitor replication lag and reconcile data; “CDC enabled” is not proof of zero data loss.
- RDS/Aurora paths include compatible dump/restore, snapshots, replicas and DMS. Check engine/version support and downtime tolerance before selecting a path.
- VM Import/Export moves supported VM images. Application Migration Service (MGN), now documented as AWS Transform MGN, replicates servers for rehosting and cutover, with staging resources and costs. Rehosting differs from moving only a database.
- Application Discovery Service is a historical inventory/discovery tool closed to new customers, not a universal replacement for MGN. VMware Cloud on AWS preserves VMware-based operating assumptions but requires current commercial availability and capacity checks. It remains named in the published exam list; that does not authorize a new subscription in this lab.
- For large datasets, estimate effective transfer time from data size and throughput, including validation and changes during transfer. Compare DataSync, suitable uploads and supported alternatives; transfer/storage notes cover the tool boundaries. Snow Family remains a published exam concept despite lifecycle changes; no physical job is created here. Consult service status.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Restore rarely, tolerate substantial recovery work | Backup and restore |
| Keep critical data/core services ready but provision serving tiers later | Pilot light |
| Need a functional secondary environment before failure | Warm standby |
| Minimize database migration downtime | Full load plus CDC, validation and controlled cutover |
| Rehost whole servers with limited rewriting | MGN |
| Convert between database engine families | Schema assessment/conversion plus a compatible data migration |
| Centralize protection across supported resources | AWS Backup plans, selections and restore tests |
| A standby cannot scale during disaster | Inspect quotas, capacity, address space and dependencies |
Traps
- Faster compute during restore improves no data that is absent from the backup.
- A replica is not automatically a historical backup, and a backup is not automatically immediately available serving capacity.
- Retention locks deliberately constrain deletion. Terraform lifecycle flags cannot bypass a genuine compliance retention requirement.
27 · More Solution Architectures
Memory hook: Choose the delivery, state and failure behavior first, then select services that preserve those decisions.
Must remember
Buffer, broadcast or replay
- A queue buffers work for competing consumers; a topic fans out events to subscribers; a retained stream supports replay and independent consumer positions. If two systems must each process every order, one shared work queue is the wrong distribution model; fan out into separate queues or use a suitable retained-stream design.
- At-least-once delivery needs idempotency. Use a stable business identifier and a durable mechanism such as a suitable conditional write or an external provider's idempotency key. Remember the uncertain outcome: a payment might succeed before a worker crashes without acknowledging the message.
- Visibility timeout gives a worker time to finish before another consumer can receive the message again; it does not prove exactly-once execution. Tune it with function duration, batching, retry and failure handling. Dead-letter handling needs an investigation/redrive process, not just a queue that silently accumulates failures.
- Scale consumers using a metric related to work and processing capacity, such as backlog per worker or message age, rather than assuming CPU always represents queue pressure. Downstream database limits can still constrain a rapidly growing worker fleet.
- EventBridge is useful for filtering/routing application and service events; Step Functions coordinates workflows with state, choices, retries and waits. Routing an event and tracking a business process are separate responsibilities. See messaging and serverless.
Cache the correct thing
- Client caches avoid a request; edge caches avoid a distant origin fetch; application/database caches avoid repeated computation or database reads. Put reusable work near its consumer while preserving security and freshness requirements.
- Cache-aside/lazy loading: the application reads the cache, fetches from the source on a miss and populates it. Write-through: update the cache with the write path. Neither avoids the need to design failure behavior and a source of truth.
- TTL and invalidation control staleness. A long TTL improves hit rates but can serve obsolete data; invalidation and immutable versioned object names solve different update patterns. A cache must not leak one user's personalized response to another because the key omitted an authorization-relevant attribute.
- A thundering herd of misses can overload the origin. Consider bounded retries, suitable cache population and origin capacity; adding a cache does not eliminate the need for a workable miss path. Edge copies are not durable backups.
Block traffic at the relevant layer
- NACLs can deny packets at a subnet boundary; security groups allow traffic but do not provide explicit deny rules. WAF IP sets and rules can filter supported web endpoints, including CloudFront. Geographic restrictions match countries, not an arbitrary list of individual IPs.
- An edge-proxied request does not necessarily expose the original client IP as the packet source at the origin. Place controls where the relevant identity/address is available and use supported forwarded-address handling carefully.
HPC and single-instance recovery
- ENA supplies enhanced networking. EFA supports compatible low-latency HPC/ML communications and OS-bypass capabilities. Cluster placement groups improve proximity for supported instance types; they do not provide multi-AZ resilience.
- FSx for Lustre provides parallel filesystem access. Batch schedules jobs and manages compute environments. ParallelCluster helps deploy HPC clusters and scheduler integrations. A scheduler, an interconnect and shared storage solve different parts of the job.
- Independent fault-tolerant batch jobs may suit Spot; tightly coordinated work needs interruption-aware planning and suitable capacity.
- An ASG with desired capacity one can replace a failed instance, including in another configured AZ, but it has a service gap. Externalize durable state and account for address changes. EBS is AZ-scoped, so a volume cannot simply be attached to a replacement in another AZ.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Survive producer bursts while workers catch up | Queue and controlled consumer scaling |
| Independent systems each require every event | Fan-out or independent stream consumers |
| Reprocess yesterday's retained telemetry | Stream with suitable retention and positions |
| Coordinate several retries and branching steps | Step Functions |
| Repeated immutable global downloads | S3 origin and CloudFront cache |
| Block abusive HTTP clients at the edge | WAF on the supported distribution |
| Tightly coupled distributed computation | EFA, cluster placement and compatible software |
| Large shared parallel-file workload | FSx for Lustre |
Traps
- Longer visibility and successful execution do not eliminate every duplicate side effect.
- Cluster placement improves proximity, while spreading capacity across failure domains improves resilience; these goals can conflict.
- Self-healing a single instance is not the same as having simultaneous healthy capacity during failure.
28 · Other Services
Memory hook: Separate deployment, administration, cost evidence and application delivery before choosing a tool.
Must remember
Ownership and controlled administration
- CloudFormation manages stack resources. Its logical resource ID, the AWS physical ID and a Terraform wrapper's resource address are different identities. Importing a stack-owned child into another state without transferring ownership creates competing lifecycle control.
- A CloudFormation service role supplies execution permissions. Associating or changing it requires appropriate iam:PassRole permission; users authorized to operate an existing stack can use its already-associated role without their own PassRole permission. Scope both the role and access to the stack carefully.
- Systems Manager Session Manager provides interactive access without inbound SSH when the managed agent, IAM and service connectivity are configured. Run Command executes documents across managed nodes. Patch Manager applies approved patch baselines; Automation orchestrates multi-step operational runbooks. A document's existence does not execute it.
- Service Catalog distributes approved infrastructure products with governance constraints, useful when teams should self-provision standardized options. License Manager helps track and manage software-license usage and rules; it does not magically supply licenses or eliminate vendor contract requirements.
- The AWS CLI and Management Console both require appropriate identities and permissions; neither bypasses IAM.
Match the money question
- Pricing Calculator: estimate a proposed architecture before deployment. Include storage, requests, transfer and networking rather than only compute.
- Cost Explorer: explore historical cost/usage, trends and forecasts. Cost and Usage Reports (CUR) provide detailed billing data for custom analysis, commonly delivered to S3; current Data Exports options extend the reporting choices.
- AWS Budgets: compare spend or usage against thresholds and notify or run separately configured actions. Cost Anomaly Detection: identify unusual spending patterns. Neither is a guaranteed immediate account spending cap; billing and notification delays matter.
- Compute Optimizer: suggests resource configuration improvements using observed utilization for supported resources. Validate recommendations against peak demand, application performance and recovery capacity; a quiet monitoring window can hide important workload needs.
- Cost allocation tags help attribute spend after required billing activation; they do not guarantee complete historical attribution. Consolidated billing changes cost visibility and potential discount sharing, not resource ownership.
- Right-size and remove waste before evaluating commitments. Stable workloads may justify Reserved Instances or Savings Plans; interruptible work may suit Spot. Compare operational effort and commitment risk with the compute choices in capacity notes.
- Instance Scheduler is a deployed AWS solution for scheduled resource stops/starts. A stopped instance can still leave storage, snapshots and some address/network charges. Scheduling is not complete teardown.
Delivery, media and application integration
- SES sends email with identity verification and applicable sending restrictions. Pinpoint engagement is the historical campaign/journey service: closed to new customers, with support scheduled to end October 30, 2026, still upcoming on this pack's October 9 review. Channel APIs have separate migration paths.
- Elastic Transcoder retired November 13, 2025 but remains named in the published exam service list. Retain the media-transcoding concept; evaluate a supported service such as MediaConvert for real deployments. Kinesis Video Streams ingests and manages video streams for playback/processing; it is different from generic ordered application records in Kinesis Data Streams.
- Batch queues jobs and orchestrates compute environments; underlying compute/storage still bill. AppFlow transfers data between supported SaaS and AWS systems; connectors and external authorization determine feasibility.
- Amplify provides frontend/application development and delivery tooling, with build, hosting and backend resource lifecycles. Device Farm tests applications on supported device/browser environments; it is a testing service, not the application's production hosting platform.
- Check service status: exam scope and deployment availability differ. Legacy exam concepts never justify prohibited subscriptions or costly capacity here.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Estimate a design before launch | Pricing Calculator |
| Explore monthly spend trends | Cost Explorer |
| Analyze detailed billing line items yourself | CUR/Data Exports and an analytical pipeline |
| Recommend resource sizing from observed utilization | Compute Optimizer |
| Publish approved self-service infrastructure options | Service Catalog |
| Track supported software-license usage | License Manager |
| Apply approved OS patches across a fleet | Patch Manager |
| Move records from a supported SaaS system to S3 | AppFlow |
| Test a mobile app across real devices | Device Farm |
Traps
- A recommendation is not automatically a safe production change; validate business and recovery requirements.
- Setting a budget does not instantly stop every charge, and stopping compute does not delete its other resources.
- An existing stack's service-role permissions can exceed a caller's direct permissions; control who may operate that stack.
29 · Whitepapers & Architectures
Memory hook: Start with the workload's requirements, then explain what the design protects, what it costs and how you know it works.
Must remember
Six pillars, six different questions
- Operational excellence: can the team operate and improve the workload? Prefer repeatable deployments, observability, runbooks, incident learning and clear response ownership.
- Security: are identities, data and workload boundaries protected? Use least privilege, appropriate encryption, traceability and a response plan. Security includes application behavior and the handling of credentials, not only network rules.
- Reliability: can the workload withstand and recover from failures while meeting its objectives? Remove unjustified single points of failure, test recovery, plan quotas/capacity and automate suitable repair. Durable data and continuously available service are different outcomes.
- Performance efficiency: do resources suit changing demand? Measure bottlenecks and choose appropriate data models, compute, storage, networking and scaling; larger instances do not fix every bottleneck.
- Cost optimization: does spending create useful value? Evaluate utilization, commitments, transfer, managed-service tradeoffs and operating effort; the lowest resource price can hide wider costs.
- Sustainability: can the workload deliver useful work with a smaller resource footprint? Improve utilization, reduce unnecessary processing/data movement and match capacity to demand. Deleting useful redundancy without considering the reliability objective is not a complete optimization argument.
- Operate, protect, recover, perform, justify cost, reduce waste: explain when a choice helps one pillar but constrains another.
Shared responsibility is service-dependent
- AWS is responsible for security of the cloud; customers remain responsible for their configuration and use of it. The boundary shifts with the service model, not with whether a console calls it “managed.”
- With EC2, the customer manages the guest OS, patching and application configuration alongside data and identities. A managed database offloads substantial engine/infrastructure work, but the customer still chooses access, data protection settings and application behavior within the service's capabilities.
- With S3, AWS operates storage infrastructure; customers control data classification, authorization and protection/retention settings. Encryption does not fix overly broad authorized access.
- Map compliance requirements to controls and evidence. AWS reports do not certify your application; see security notes.
Review tools support judgment
- Well-Architected Tool records reviews, answers, risks and milestones. Assign owners and prioritize evidence-backed remediation; completing a questionnaire does not automatically make the workload resilient or compliant.
- Trusted Advisor supplies checks and recommendations, with availability varying by support/features. A recommendation is an investigation lead, not permission to change production. Review workload purpose, peak demand and recovery constraints before accepting a cost or capacity suggestion.
- Reference architectures and AWS Architecture Center examples illustrate patterns, service boundaries and failure domains. Adapt them to access patterns, compliance, team skills, latency, RPO/RTO and budget. A three-Region diagram is not a requirement for every application.
- Specify traffic, consistency, tolerable data loss, recovery time and failover capacity. Replace “multi-AZ therefore highly available” with concrete failure behavior and tested recovery.
Read a design question systematically
- Extract hard constraints first: residency, compatibility, loss/outage tolerance, access pattern and operational requirements. Eliminate options that violate them before comparing cost or convenience.
- Match the stated optimization: operational effort, resilience, cost or latency can favor different solutions. Avoid unneeded capabilities.
- Distinguish the four exam domains from the six pillars. The published SAA-C03 outline organizes assessment around secure, resilient, high-performing and cost-optimized architecture; the framework is the broader review lens. The official guide explicitly says its content list is non-exhaustive, so no notes pack can guarantee every possible question.
- Validate data and failure paths: nodes, AZs, dependencies, credentials and operator mistakes. Include teardown and ownership for disposable environments.
Choose under exam pressure
| Clue in the requirement | Decision habit |
|---|---|
| Fragile single component threatens the availability target | Review reliability and failure domains |
| Repeated manual recovery steps create mistakes | Improve operational automation and runbooks |
| Idle resources consume budget and power | Assess utilization, cost and sustainability together |
| Required data residence conflicts with a cheaper Region | Satisfy the hard constraint before optimizing price |
| A reference design has more Regions than needed | Reassess against actual RPO/RTO and operational complexity |
| A recommendation conflicts with failover capacity | Validate workload purpose before applying it |
| A managed service stores sensitive customer data | Identify the customer's remaining access/configuration duties |
Traps
- Two subnet AZs do not prove there are two simultaneously healthy application targets.
- A passing configuration check does not measure real recovery, application correctness or compliance.
- “AWS manages the service” does not remove the customer's data, identity and application responsibilities.
30 · Organisational Architecture and Guardrails
Memory hook: Design the account, identity, network and evidence boundaries before centralising services.
Must remember
- Use accounts as isolation and ownership boundaries. Organise OUs by control requirements rather than copying every management-chart level. Separate security/logging, shared services, networking and workloads where justified; avoid routine workloads in the organisation management account.
- Identity Center/federation gives workforce access through scoped roles. SCPs limit applicable permissions but do not grant them. Resource policies, organisation conditions, permission boundaries and delegated administration solve different cross-account problems. Design and test emergency access rather than removing controls during an outage.
- Centralise logging and evidence with protected destinations, KMS policies and retention. Control Tower can establish a governed landing zone; account vending and StackSets automate consistent baselines. Existing-account onboarding still requires conflict, ownership and remediation planning.
- Choose peering, Transit Gateway, Cloud WAN or service-level PrivateLink according to topology, segmentation, routing scale and operational ownership. Centralised egress/inspection creates dependencies and transfer costs; verify symmetric stateful paths and per-AZ resilience. Hybrid DNS requires Resolver endpoints/rules and zone association, not merely network connectivity.
- Share supported resources through RAM where appropriate instead of duplicating everything or granting broad cross-account administration. Check quotas, Region boundaries, service compatibility and the failure domain of each central service.
- Cost visibility needs activated allocation tags, linked-account ownership, exports and budgets. Purchase commitments only after rightsizing and measuring stable usage. Chargeback/showback should distinguish shared platform costs from workload-controlled spending.
- Governance has preventive, detective and corrective controls. Apply proportionate guardrails, document exceptions, test changes in a limited OU and preserve business continuity. A centrally enforced error can have a larger blast radius than a local error.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Many accounts need consistent infrastructure baselines | Account provisioning plus StackSets/pipelines and scoped roles. |
| Teams need one service without broad network access | PrivateLink where the service pattern fits. |
| Security needs tamper-resistant organisation evidence | Centralised protected logs, encryption permissions and appropriate retention. |
Traps
- Consolidated billing does not merge account permissions.
- One central firewall can become a shared failure/cost bottleneck.
- An SCP applied broadly can block required service operations if exceptions are poorly designed.
31 · Professional Architecture Decisions and Migration Sequencing
Memory hook: Reject any option that breaks a hard constraint before comparing convenience or price.
Must remember
- Extract availability, RPO/RTO, residency, latency, throughput, consistency, licensing, skills and cost constraints. Separate mandatory requirements from preferences. Prefer the least complex design that satisfies all mandatory constraints, not the service with the most features.
- Compare backup/restore, pilot light, warm standby and active-active using recovery time, data loss, capacity and operating complexity. Database replication is commonly asynchronous across Regions; measure lag and define promotion, write ownership and failback. Active-active writes add conflict and consistency problems.
- Map application dependencies and select migration waves. Rehost, replatform, refactor, repurchase, relocate, retain and retire have different effort and business outcomes. Application Migration Service, DMS/CDC, DataSync and transfer services move different assets; schema and application compatibility remain separate.
- Plan initial load, ongoing change capture, validation, cutover window, rollback and source retirement. A database cutover needs a strategy for writes occurring after the switch; simply pointing DNS back can lose or fork data. TTL/cache behaviour influences traffic transition.
- Modernise selectively: externalise state, decouple with queues/events, use managed data services and replace only the components whose constraints justify it. Strangler-style migration routes selected functionality to a new implementation while preserving a controlled path back.
- Improve existing systems using evidence: traces/metrics, query plans, right-sizing, caching, partitioning, autoscaling, storage lifecycle and network transfer analysis. A cache changes freshness; read replicas do not solve every write bottleneck; extra compute cannot fix lock contention.
- Compare complete solution cost, including licences, transfer, minimum storage duration, requests, support and staff effort. Test availability and recovery claims with meaningful fault/restore exercises. Record assumptions and the condition that would make the chosen design need revision.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Strict low RPO with regional disaster recovery | Choose and test a replication/recovery design that meets measured loss limits. |
| Migration deadline is short and architecture change is not required | Consider rehost/replatform before a large refactor. |
| Legacy and modern components must coexist | Incremental migration with explicit routing and data ownership. |
Traps
- Lowest operational effort is a constraint only when the question asks for it.
- A DNS failback cannot reconcile divergent writes.
- An available replica is not necessarily ready for the full production load.
32 · Hybrid Networks, Shared Inspection and DNS Boundaries
Memory hook: Trace both directions through every routing and trust boundary.
Must remember
For a professional scenario, draw the complete client-to-service path and its return. Include VPC subnet routes, Transit Gateway tables, on-premises BGP, security groups/NACLs, inspection appliances and DNS. A working outbound route does not prove return reachability or stateful-session symmetry.
Transit Gateway association selects the route table used for traffic entering from an attachment; propagation advertises that attachment’s routes into selected tables. An attachment associates with one table and can propagate into several. Segment production, nonproduction and shared services by deliberate table design. A propagated prefix does not automatically create the necessary VPC subnet route. Ordinary VPC peering is not transitive, and overlapping address space prevents normal unambiguous routing.
Centralized inspection needs the return path through the stateful appliance handling the forward flow. Transit Gateway appliance mode supports flow/AZ affinity on the inspection attachment; it does not repair every incorrectly configured VPC route. Use resilient per-AZ appliance/endpoints and inspect the failure behavior. Gateway Load Balancer distributes traffic to supported virtual appliances; Network Firewall provides managed firewall capabilities. Choose from required inspection functions, operational ownership and supported routing patterns.
Direct Connect private virtual interfaces connect supported private VPC paths; transit virtual interfaces with a Direct Connect gateway support Transit Gateway connectivity. Public virtual interfaces reach AWS public service prefixes and are not ordinary internet transit. Direct Connect is not inherently an encrypted end-to-end tunnel: use appropriate TLS, VPN or supported MACsec according to the requirement. Two circuits at one location can share a location failure; diversify devices, connections and locations when resilience requires it. A VPN backup must have sufficient capacity and correct BGP preference to be useful.
PrivateLink exposes a supported service through consumer endpoints instead of granting broad routed access, which can help across overlapping consumer/provider networks. It is not a general transitive network. Interface endpoints have DNS and per-AZ cost/availability choices; gateway endpoints serve supported S3/DynamoDB VPC route-table use cases and are not universal on-premises access paths.
For hybrid DNS, inbound Resolver endpoints accept queries from external resolvers; outbound endpoints/rules forward selected queries outward. Associate private hosted zones and shared rules with the intended VPCs. Avoid forwarding a zone back to the resolver that originally forwarded it. A private hosted zone can mask public names in its namespace; missing private records do not necessarily fall back to the public zone.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Many networks require segmented routed connectivity | Transit Gateway/appropriate global networking with explicit tables. |
| Consumers need one service, including overlapping CIDRs | PrivateLink where its service pattern fits. |
| On-premises must resolve a private hosted zone | Inbound Resolver path, zone association and permitted DNS traffic. |
Traps
- Association and propagation solve different routing steps.
- A backup network path is not useful if it cannot carry recovery traffic.
- Stateful inspection needs forward and return symmetry.
33 · Cross-Account Authorization and Evidence Architecture
Memory hook: Follow the principal, the policy ceiling, the resource and the key.
Must remember
An assume-role design has two stages: the source must be allowed to request the role and the target trust policy must trust the appropriate principal/conditions; the resulting session then acts with the target role’s effective permissions. A resource-based grant can authorize supported direct cross-account access, but policy evaluation differs from assuming a role. Resource policies do not exist for every service.
Identity policies grant actions; permission boundaries and SCPs constrain applicable permissions. SCPs do not grant permissions, do not govern the management account in the same way as member accounts, and have documented exclusions such as service-linked roles. Explicit deny prevails where applicable. An allow-list SCP strategy requires appropriate allowance through the hierarchy, so a lower-level allow cannot repair an upper-level omission. Test policy changes in a limited OU with service dependencies and emergency access in view.
For third-party role access, an external ID helps address confused-deputy risk; it is not a secret password or a replacement for a precise trust principal. Service integrations may need source-account/source-ARN conditions. Organization conditions reduce repeated account lists but must match the actual supported request context; an absent condition key can change policy behavior.
Encrypted cross-account data needs both data-service authorization and the required KMS authorization. Sharing an encrypted snapshot or object while omitting key permissions is incomplete. AWS managed keys have sharing limitations that can require a copy encrypted under a customer managed key. Separate key administration from data use, and plan key retention for backups. A seven-day deletion window is still a future point of irreversible loss if the key is needed by retained data.
Central evidence should survive a workload-account compromise: organization trails where appropriate, narrowly scoped data events, protected log destinations, controlled key access and separate administration. CloudTrail records API activity; Config evaluates/configuration-records supported resources; CloudWatch observes metrics/logs. Delegated administration permits service-specific centralized operations without routine use of the management account. Preserve the ability to identify which workload principal performed the original action.
Scenario drill: an organization-wide deployment fails only when writing encrypted logs. Check the executing identity, service trust, SCP/boundary, destination policy and KMS path before granting administrator permissions. Diagnose the failed authorization edge rather than widening every boundary.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| External vendor administers a limited account function | Narrow cross-account role, precise trust and external ID as appropriate. |
| Object readable by policy but KMS denies decryption | Repair the key authorization path, not just the bucket policy. |
| Compromised workload must not erase its own evidence | Separately administered protected logging destinations. |
Traps
- External ID is not authentication by itself.
- An SCP allow does not grant an IAM action.
- Access to encrypted data can fail at either the resource or key boundary.
34 · Regional Recovery, Write Ownership and Failback
Memory hook: Traffic can switch faster than data can become safe.
Must remember
Set RPO and RTO for the business operation, then examine every dependency: database, object store, identity, keys, secrets, DNS, images, quotas and deployment artifacts. Multi-AZ availability does not alone provide regional disaster recovery. A standby application with no accessible encryption key or sufficient quota is not ready.
Backup/restore trades lower idle cost for restoration/provisioning time. Pilot light keeps critical state or core capability available while rebuilding/scaling other components. Warm standby maintains a functional reduced-capacity environment. Active-active serves from multiple locations but adds data ownership, conflict and operational complexity. The architecture name is less important than measured recovery behavior.
Aurora Global Database commonly uses cross-Region replication to a secondary Region. A planned switchover coordinates a supported controlled role change; an unplanned failover must account for replication lag and possible data loss. The application must reconnect to the intended writer and avoid sending writes to an old or isolated primary. DynamoDB global-table consistency mode and regional topology affect guarantees; do not assume every configuration provides zero-loss synchronous semantics.
S3 replication, database replicas and synchronized files can reproduce unwanted changes. Retained backups/version history address historical recovery, subject to replication configuration, retention and keys. Cross-account backup isolation protects against a different threat from regional redundancy. Verify supported resource/copy/encryption combinations instead of assuming every AWS Backup feature applies to every service.
Route 53 health-based routing and ARC routing controls can assist controlled traffic movement for appropriate designs. DNS caching means existing clients may not move immediately. Global Accelerator can provide a stable anycast entry point with endpoint health routing, but cannot repair inconsistent application data. ARC zonal shift addresses an AZ impairment, not a complete database cross-Region promotion plan.
Failback is another migration. Reconcile writes, establish replication in the new direction where supported, verify consistency, then switch ownership/traffic in a controlled window. Simply returning DNS to the original Region can create split-brain or discard newer transactions. Run recovery exercises that measure user-visible restoration and data correctness, not merely whether infrastructure creation succeeded.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Cheap recovery with hours of acceptable outage | A tested backup/restore design may fit. |
| Minutes of recovery with reduced standby capacity | Warm standby with validated scaling and dependencies. |
| Fail back after writes in the recovery Region | Reconcile/synchronize data and control writer ownership before switching traffic. |
Traps
- Healthy endpoints do not prove consistent data.
- Asynchronous replication lag matters to RPO.
- A DNS change is not a database conflict-resolution mechanism.
35 · Migration Waves, Database Cutover and Strangler Patterns
Memory hook: Discover together, move deliberately, reconcile before retirement.
Must remember
Inventory business services and their technical dependencies, including batch jobs, shared files, identity, licensing, network allow lists and hidden database consumers. Group tightly coupled components into migration waves or provide a temporary connectivity strategy. Prioritize using business value, risk, readiness and dependency complexity, not only server size.
Choose among retain, retire, rehost, relocate, repurchase, replatform and refactor. A deadline-driven data-center exit may favor rehost before selective modernization; a license or engine limitation may require a different path. AWS Application Migration Service supports server rehosting; database migration needs its own schema/application/data plan. Service availability and source support must be checked for the actual environment.
For databases, separate schema conversion from data movement. DMS supports full load and CDC for supported engines; conversion tooling addresses compatible schema/code transformations but cannot guarantee every stored procedure or application query behaves identically. Assess data types, encoding, constraints, sequences, large objects and engine-specific features. Validate counts and business totals as well as transport status.
Cutover sequence: prepare target and recovery plan, load baseline, capture changes, reconcile, control source writes, drain the replication gap, switch clients, validate business behavior and monitor. Rollback must account for writes accepted at the new target. Keep the source only as long as its defined rollback/compliance role requires, then retire credentials, replication jobs and temporary connectivity.
DataSync moves supported file/object data and helps recurring transfer; Transfer Family exposes managed transfer protocols; Storage Gateway supports ongoing hybrid access patterns. Pick by application protocol and operating requirement. A physical-transfer solution may be inappropriate or unavailable to new customers; verify current eligibility and lead time rather than choosing it solely because a dataset is large.
A strangler approach moves selected capabilities behind controlled routing while the legacy system remains. Use an anti-corruption layer to isolate old data/contracts, decide authoritative ownership and reconcile events. The transactional outbox pattern can reduce the risk of committing database state without its corresponding event; consumers still need idempotence. A distributed saga uses local transactions and compensating actions where a single global transaction is unsuitable. Compensation is a business action, not a time machine that undoes all side effects.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Short exit deadline with compatible servers | Rehost with dependency-aware waves and later optimization. |
| Low-downtime supported DB migration | Full load plus CDC, reconciliation and controlled cutover. |
| Replace one legacy capability at a time | Strangler routing with explicit data ownership. |
Traps
- CDC is not schema conversion.
- Dual writes can diverge when one destination succeeds and the other fails.
- A migration tool’s success status is not business acceptance.
36 · Performance Evidence, Cost Allocation and Commitment Risk
Memory hook: Measure the bottleneck; price the whole path; commit only to the baseline.
Must remember
Translate performance objectives into measurable latency percentiles, throughput, concurrency and freshness. Average latency can hide severe tail delays. Correlate traces, service metrics, database query plans and resource saturation to locate the limiting component. Scaling a web tier will not fix a serialized database lock or a single hot partition.
Caching moves reads away from slower dependencies but changes freshness, invalidation and failure behavior. Select TTL and eviction based on data semantics; protect against cache stampedes and cold-cache load. Read replicas distribute eligible reads, while connection pooling/proxying addresses connection pressure. Neither automatically increases the writer’s transaction capacity. Sharding/partition-key changes can improve distribution but introduce query and operational trade-offs.
Match storage performance to access pattern, not just capacity. IOPS, throughput, latency and request size interact. A migration from an expensive disk class to a cheaper one is acceptable only if required sustained/burst behavior remains supported. Small-file/request-heavy object access may be dominated by request/transition costs rather than bytes stored.
Cost allocation starts with accounts, activated cost-allocation tags and consistent ownership. Cost and Usage Report/Data Exports support detailed analysis; Cost Explorer supports exploration; budgets and anomaly detection notify. They do not universally prevent spend. Allocate shared networking, security and platform costs using an agreed showback/chargeback method rather than treating untagged costs as free.
Rightsize, schedule idle environments and remove unnecessary retained resources before making commitments. Savings Plans exchange eligible spend commitment for rates; Reserved Instances have offering-specific discount/capacity attributes; Spot trades interruption/capacity risk for lower price. Do not confuse a billing discount with a guarantee that a required AZ/type is available. Evaluate utilization and coverage separately, and preserve flexibility for uncertain migration demand.
Compare full data paths: NAT processing, endpoint hourly/data charges, cross-AZ/Region transfer, replication, log ingestion/retention, requests, licenses and operator effort. Centralizing egress can reduce duplicated infrastructure while increasing transit/transfer or failure dependencies. Choose the cheapest design that satisfies the full requirement, then validate assumptions with actual usage.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Database high latency despite idle app CPU | Inspect query plans, locks, connections and storage before scaling web servers. |
| Stable measured compute baseline | Evaluate an appropriate commitment after rightsizing. |
| Large NAT bill from supported AWS-service traffic | Evaluate endpoint routing and its complete cost/availability trade-offs. |
Traps
- Lowest hourly instance price is not lowest solution cost.
- A cache can worsen recovery if every instance misses simultaneously.
- Commitment coverage and commitment utilization answer different questions.