Memory hook: Raw to trusted, with replayable progress.
Must remember
- Use bronze/silver/gold as a useful layering pattern: retain raw evidence, validate/conform it, then publish business-ready data. Choose grain and SCD behavior from the required history.
- Delta adds transaction-log semantics to data files; Parquet is a columnar format; CSV/JSON are interchange formats; Iceberg is another table format with its own support considerations. Choose compatibility intentionally.
- Lakeflow Connect, notebooks, Data Factory, SQL COPY INTO/CTAS and supported connectors solve different ingestion needs. Batch suits bounded arrivals; Structured Streaming suits incremental unbounded processing.
- Auto Loader incrementally discovers supported files with checkpoints/schema handling. Event Hubs can feed streaming pipelines through supported interfaces; configure authentication, offsets and consumer behavior.
- CDC and MERGE support incremental updates when keys and ordering are correct. Deduplicate source changes and handle late/out-of-order records so one key is not matched ambiguously.
- Partitioning, Z-ordering and liquid clustering optimize different layouts; use supported combinations rather than piling every technique onto a small table. Managed versus external storage determines lifecycle ownership.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Files continuously arrive in cloud storage | Auto Loader with durable checkpoints and deliberate schema evolution. |
| Need to preserve changing dimension history | SCD Type 2 with correct effective periods and change ordering. |
Traps
- Streaming is not automatically lower cost than a frequent batch job.
- Partitioning by a very high-cardinality field can create excessive small files.
Active recall
1. What does a checkpoint preserve?
Streaming progress/state needed for supported recovery and continued processing.
2. How does SCD Type 1 differ from Type 2?
Type 1 overwrites current attributes; Type 2 retains historical versions.
3. Why deduplicate before MERGE?
Multiple source rows for one target key can create ambiguity or incorrect updates.
4. What does CTAS do?
Creates a table from a query result.
5. Why retain a raw layer?
To support audit, replay and correction when transformation logic changes.