Memory hook: Version the job; inspect the expensive stage.
Must remember
- Use Git branches, reviews and tests for notebooks/code/configuration. Unit, integration, end-to-end and user-acceptance tests validate different layers; include data correctness and access boundaries.
- Declarative Automation Bundles package supported jobs/pipelines and environment targets for repeatable deployment. Validate, deploy and run through supported CLI/API workflows with scoped service identities.
- Inspect Spark UI, DAG/stage metrics and query profiles to distinguish shuffle, skew, spill, cache misses and insufficient parallelism. One slow skewed partition can dominate an otherwise idle cluster.
- Tune joins, data layout, partition count and compute only after measuring. Broadcast joins suit genuinely small inputs; excessive caching can consume memory and increase spill.
- OPTIMIZE compacts/reorganizes supported Delta layouts; VACUUM removes obsolete files according to retention and safety constraints. Aggressive retention can break time travel or running readers.
- Monitor cluster consumption, job duration, failures and cost. Use supported Log Analytics integration and Azure Monitor alerts; repair/restart jobs only with an understanding of checkpoints and side effects.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| One Spark task runs much longer than others | Investigate skew and partition distribution. |
| Many tiny Delta files slow reads | Evaluate OPTIMIZE and better write patterns before adding nodes. |
Traps
- VACUUM is not a generic performance command with no recovery consequences.
- Restarting a cluster does not fix a deterministic bad join or data-quality error.
Active recall
1. What does spilling indicate?
Intermediate data exceeded available execution memory and was written to disk.
2. Why deploy a bundle per environment?
To keep reviewed definitions consistent while varying approved configuration.
3. How verify a performance change?
Compare stage metrics, duration, cost and output correctness on representative data.
4. What can low retention remove?
Files needed by historical versions or concurrent readers, depending on timing and support.
5. When is a repair run appropriate?
When the failed portion can safely resume/re-execute with consistent inputs and side effects.