certslothcertsloth
SAA-C03/Topic 27

AWS / Associate

More Solution Architectures

6 min read5 recall promptsReviewed 2026-10-10

Memory hook: Choose the delivery, state and failure behavior first, then select services that preserve those decisions.

Must remember

Buffer, broadcast or replay

  • A queue buffers work for competing consumers; a topic fans out events to subscribers; a retained stream supports replay and independent consumer positions. If two systems must each process every order, one shared work queue is the wrong distribution model; fan out into separate queues or use a suitable retained-stream design.
  • At-least-once delivery needs idempotency. Use a stable business identifier and a durable mechanism such as a suitable conditional write or an external provider's idempotency key. Remember the uncertain outcome: a payment might succeed before a worker crashes without acknowledging the message.
  • Visibility timeout gives a worker time to finish before another consumer can receive the message again; it does not prove exactly-once execution. Tune it with function duration, batching, retry and failure handling. Dead-letter handling needs an investigation/redrive process, not just a queue that silently accumulates failures.
  • Scale consumers using a metric related to work and processing capacity, such as backlog per worker or message age, rather than assuming CPU always represents queue pressure. Downstream database limits can still constrain a rapidly growing worker fleet.
  • EventBridge is useful for filtering/routing application and service events; Step Functions coordinates workflows with state, choices, retries and waits. Routing an event and tracking a business process are separate responsibilities. See messaging and serverless.

Cache the correct thing

  • Client caches avoid a request; edge caches avoid a distant origin fetch; application/database caches avoid repeated computation or database reads. Put reusable work near its consumer while preserving security and freshness requirements.
  • Cache-aside/lazy loading: the application reads the cache, fetches from the source on a miss and populates it. Write-through: update the cache with the write path. Neither avoids the need to design failure behavior and a source of truth.
  • TTL and invalidation control staleness. A long TTL improves hit rates but can serve obsolete data; invalidation and immutable versioned object names solve different update patterns. A cache must not leak one user's personalized response to another because the key omitted an authorization-relevant attribute.
  • A thundering herd of misses can overload the origin. Consider bounded retries, suitable cache population and origin capacity; adding a cache does not eliminate the need for a workable miss path. Edge copies are not durable backups.

Block traffic at the relevant layer

  • NACLs can deny packets at a subnet boundary; security groups allow traffic but do not provide explicit deny rules. WAF IP sets and rules can filter supported web endpoints, including CloudFront. Geographic restrictions match countries, not an arbitrary list of individual IPs.
  • An edge-proxied request does not necessarily expose the original client IP as the packet source at the origin. Place controls where the relevant identity/address is available and use supported forwarded-address handling carefully.

HPC and single-instance recovery

  • ENA supplies enhanced networking. EFA supports compatible low-latency HPC/ML communications and OS-bypass capabilities. Cluster placement groups improve proximity for supported instance types; they do not provide multi-AZ resilience.
  • FSx for Lustre provides parallel filesystem access. Batch schedules jobs and manages compute environments. ParallelCluster helps deploy HPC clusters and scheduler integrations. A scheduler, an interconnect and shared storage solve different parts of the job.
  • Independent fault-tolerant batch jobs may suit Spot; tightly coordinated work needs interruption-aware planning and suitable capacity.
  • An ASG with desired capacity one can replace a failed instance, including in another configured AZ, but it has a service gap. Externalize durable state and account for address changes. EBS is AZ-scoped, so a volume cannot simply be attached to a replacement in another AZ.

Choose under exam pressure

Clue in the requirement Choose or investigate
Survive producer bursts while workers catch up Queue and controlled consumer scaling
Independent systems each require every event Fan-out or independent stream consumers
Reprocess yesterday's retained telemetry Stream with suitable retention and positions
Coordinate several retries and branching steps Step Functions
Repeated immutable global downloads S3 origin and CloudFront cache
Block abusive HTTP clients at the edge WAF on the supported distribution
Tightly coupled distributed computation EFA, cluster placement and compatible software
Large shared parallel-file workload FSx for Lustre

Traps

  • Longer visibility and successful execution do not eliminate every duplicate side effect.
  • Cluster placement improves proximity, while spreading capacity across failure domains improves resilience; these goals can conflict.
  • Self-healing a single instance is not the same as having simultaneous healthy capacity during failure.

Active recall

1. A payment succeeds, then its Lambda invocation times out. How can the retry avoid charging again?

Use a durable idempotency mechanism tied to the payment's stable business identifier. A timeout leaves the result uncertain; longer visibility alone does not tell the retry whether the side effect already occurred.

2. Inventory and email systems must independently receive every order. Why not let both poll one queue?

Competing consumers generally split the work. Fan-out into separate queues or suitable independent stream consumption preserves each system's ability to process every event and recover separately.

3. A CDN serves one customer's account page to another. What should be investigated besides encryption?

Investigate whether personalized responses were cacheable and whether the cache key and authorization design separated users correctly. TLS protects transport; it does not correct an unsafe cache boundary.

4. An HPC application spends most of its time communicating among nodes. Why might bigger CPUs be insufficient?

Its bottleneck may be interconnect latency or throughput. Evaluate compatible EFA networking, placement, software and shared storage rather than purchasing CPU capacity that leaves the communication delay unchanged.

5. An ASG replaces its only failed instance in another AZ. Why can the application still be unavailable?

Replacement and initialization take time, there was no second live target, and local/AZ-scoped state might not follow. Recovery requires durable external state and enough simultaneous capacity for the availability objective.

Terraform anchor: A dependency graph and timeout precondition establish configuration assumptions; neither supplies idempotent runtime processing or continuous availability.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.