certslothcertsloth
DP-700/Topic 03

Microsoft / Associate

Pipelines, Dataflows and Incremental Loads

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Orchestrate dependencies; move data; make retries safe.

Must remember

Pipelines coordinate activities such as copy, notebooks and dataflows with dependencies, parameters and dynamic expressions. Dataflows Gen2 offer Power Query transformations; notebooks support code-oriented Spark and other supported execution. Choose the tool from complexity, language, scale and maintenance needs rather than assuming one tool fits every step.

Full loads replace or reload a complete dataset; incremental loads process changes since a tracked boundary. A watermark based on a timestamp/key must handle ties, updates, deletes and clock/source semantics. CDC captures changes from a supported source; mirroring replicates supported source changes into Fabric with product-specific limits. Neither guarantees a correct dimensional model automatically.

Use schedules when cadence is time-driven and event triggers when supported events should initiate work. Pass environment and partition parameters explicitly. Model retries, dependencies, concurrency and failure paths; a partially successful run must not duplicate or silently omit records. Stage data, validate it and publish atomically where the engine supports the required semantics.

Track source positions, row counts and business reconciliation totals. Treat late-arriving dimensions/facts deliberately, preserving unknown-member or correction logic. Protect credentials in managed connections and avoid writing secrets into parameters/logs. For Airflow-based orchestration, keep heavy data processing in the appropriate engine rather than passing large datasets through orchestration metadata.

Choose under exam pressure

Requirement Choice and reason
Mostly visual cleanup and shaping Dataflow Gen2.
Complex PySpark transformation Notebook.
Coordinate copy and dependent transforms Pipeline or appropriate Airflow workflow.

Traps

  • A timestamp watermark can miss records if its boundary logic is wrong.
  • A successful copy is not proof that deletes and updates were handled correctly.

Active recall

1. Full versus incremental load?

All required data versus only changes since a tracked boundary.

2. Why make retries idempotent?

An earlier attempt may have committed part or all of its output.

3. What does mirroring provide?

Supported ongoing source replication into Fabric, with source-specific behavior.

4. Why reconcile business totals?

Technical success can conceal missing, duplicated or misinterpreted records.

5. Why parameterize partitions?

To make scheduled runs and backfills explicit and reproducible.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.