Memory hook: Quality gates before dependent jobs.
Must remember
- Profile distributions, nulls, duplicates and types before transformations. Join, union, intersect and except have different semantics; validate row counts after joins to detect accidental multiplication.
- Use filters, groups, aggregates, pivots/unpivots and denormalization according to the target grain. Choose merge/insert/append based on update semantics rather than convenience.
- Schema enforcement rejects incompatible writes; schema evolution accepts supported planned changes. Uncontrolled drift can silently change analytics meaning even when the pipeline runs.
- Lakeflow Spark Declarative Pipelines express managed transformations and expectations. Configure whether failed expectations are retained, dropped or fail processing according to supported behavior and business policy.
- Lakeflow Jobs orchestrates tasks with dependencies, parameters, triggers/schedules, retries and notifications. A notebook task and a declarative pipeline have different lifecycle and dependency management models.
- Design failure handling, checkpoints and repair strategy before production. Retrying only the failed task is safe only when upstream/downstream side effects and data versions remain consistent.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Invalid records should be inspected without blocking valid ones | A deliberate quarantine/expectation policy with visible quality metrics. |
| Several dependent transformations run nightly | Lakeflow Jobs with explicit dependencies and failure notifications. |
Traps
- A successful Spark job can still duplicate every business measure through a bad join.
- Automatically accepting every new column is not a governance strategy.
Active recall
1. Why validate cardinality after a join?
To catch unexpected one-to-many/many-to-many expansion.
2. What is a pipeline expectation?
A declarative data-quality condition with configured handling for violations.
3. How choose append versus merge?
Append adds new records; merge handles keyed changes when update semantics require it.
4. Why make tasks idempotent?
Retries and repairs should not duplicate business output.
5. What should a job alert contain?
Failed task/run, data impact, logs and the responsible recovery owner.