certslothcertsloth
MLA-C02/Topic 02

AWS / Associate

Ingestion, Streaming and Transformation

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: A durable checkpoint and an idempotent sink matter more than a promise of no retries.

Must remember

  • Choose batch, micro-batch or streaming using freshness, volume, latency and recovery needs. Full loads copy a dataset; CDC captures changes. Preserve source offsets, timestamps and identifiers so replay, deduplication and audit are possible.
  • Kinesis partition keys determine shard placement and per-key ordering; skew can overload a shard. Firehose handles supported delivery/buffering, not arbitrary consumer replay. MSK provides managed Kafka infrastructure; consumer groups and offsets have their own processing semantics.
  • Glue jobs and EMR/Spark transform data; Lambda fits bounded event processing; Managed Service for Apache Flink handles stateful streaming. Flink checkpoints preserve recoverable state; event time, processing time, watermarks and late-event policies affect window correctness.
  • In Spark, narrow transformations avoid a shuffle; wide operations such as many joins/aggregations redistribute data. Partition skew, tiny files, excessive shuffles and driver collection can dominate runtime. Broadcast a genuinely small dimension when appropriate; do not broadcast a dataset that exhausts executor memory.
  • Glue bookmarks track supported processed inputs, not universally exactly-once business output. A failed job can have written partial results. Use transactional tables, staging/commit patterns or idempotent upserts as required.
  • Step Functions, Glue workflows or managed Airflow coordinate dependencies and retries according to operational needs. Keep configuration separate from code, package dependencies reproducibly and unit-test transformations plus integration contracts. An SDK paginator is necessary when an API returns continuation tokens.

Choose under exam pressure

Requirement Choice and reason
Correct results for late-arriving event timestamps Event-time windows with a defined lateness policy.
A few keys dominate stream traffic Reconsider partitioning and downstream ordering requirements.
A join shuffles a huge fact table against a tiny dimension Evaluate a safe broadcast join.

Traps

  • Successful source reads do not prove a committed sink write.
  • A bookmark is not a substitute for business-key deduplication.
  • Increasing workers may not fix one skewed partition.

Active recall

1. Why retain source offsets?

To resume/replay deterministically and audit which input produced output.

2. What differs between event time and processing time?

When the event occurred versus when the processing system observes it.

3. Why are thousands of tiny files costly?

Listing, scheduling and per-file overhead can dominate useful scanning work.

4. What protects a sink during a retried batch?

An idempotent key/upsert, transactional commit or another explicit duplicate-safe design.

5. Why avoid collecting a large Spark dataset to the driver?

It can overwhelm driver memory and remove distributed-processing benefits.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.