Memory hook: Move once safely; transform at the right grain; observe end to end.
Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.
Scope/version: Microsoft’s public guide already describes the October 19, 2026 update, ten days after this review. This path explicitly targets that upcoming outline. For an earlier sitting, compare the outline attached to your booking; localized updates can differ.
Must remember by domain
| Domain | Rapid revision |
|---|---|
| Platform/configuration | Capacity supplies execution resources; workspace contains items; domains organize ownership/governance. Configure Spark sessions/libraries, OneLake and supported Airflow settings intentionally. Lakehouse serves Delta/Spark, Warehouse relational T-SQL, Eventhouse KQL/event analytics. |
| Lifecycle/governance | Git and deployment pipelines manage supported definitions; database projects model schemas. Environment connections, identities and data still need validation. Workspace/item/file/SQL/row/column policies govern different paths. OneLake security support depends on engine/mode. Masking changes display, not encryption; labels and endorsement do not replace access controls. |
| Orchestration | Pipelines coordinate activities; Dataflows Gen2 provide Power Query transformations; notebooks provide code-oriented processing. Schedules are time-driven; supported event triggers react to events. Pass parameters/dynamic expressions, define dependencies and bound concurrency. Airflow should orchestrate heavy processing rather than shuttle huge datasets through metadata. |
| Loading | Full load reloads the set; incremental load processes a tracked change boundary. Timestamp watermarks need ties, late updates and deletes considered. CDC/mirroring have source-specific semantics. Shortcuts reference existing data and depend on target connection/security. Reconcile row counts, business totals and checkpoints before marking success. |
| Batch transformation | Delta adds transaction metadata to lake files. Bronze/silver/gold separates raw, conformed and curated responsibility. SCD1 overwrites, SCD2 preserves history. Join at the correct grain, deduplicate by a defensible key/order, handle late dimensions/facts and avoid collecting large distributed data into a driver. |
| Streaming | Eventstream routes/transforms supported events; Eventhouse stores/queries events; Structured Streaming supports code-based state. Event time differs from processing time. Tumbling windows do not overlap; hopping/sliding can; sessions group by inactivity. Watermarks bound late-data/state handling. Native Eventhouse tables and accelerated/ordinary OneLake shortcuts have different freshness/cost behavior. |
| Monitoring/optimization | Monitor copy, transformations, table freshness, semantic refresh and consumption. Diagnose activity errors, credentials/gateway, schema, library versions and actual data paths before rerun. Spark UI exposes skew/shuffle/spill. Compaction/layout reduces small-file overhead; V-Order suits supported read optimization; VACUUM affects historical files. Measure capacity contention and downstream limits. |
Traps and recovery order
A green copy task can leave a stale dashboard. Retry only after understanding partial writes and checkpoints; use staging/validation and supported atomic publication to avoid duplicate results. Increase concurrency only when source and downstream capacity permit. Retaining source history makes replay possible, but idempotency makes replay safe.
Last-pass self-check
1. When is a timestamp-only watermark unsafe?
Ties, delayed updates, deletes or source timestamp semantics can omit or duplicate changes.
2. Which window uses a gap in activity?
A supported session window.
3. Why can a OneLake shortcut suddenly fail?
The referenced target, connection, permission or source availability changed.
4. A Spark cluster is mostly idle with one slow task: what next?
Inspect skew, partitioning and shuffle/spill in the execution stages.
5. What proves end-to-end data freshness?
Source progress, completed transformations, published tables, semantic-model state and consumer-visible results.
Sources
- Official exam scope and version
- fabric · fundamentals · microsoft-fabric-overview
- fabric · onelake · onelake-overview
- fabric · cicd · cicd-overview
- fabric · data-engineering · spark-compute
- fabric · security · security-overview
- fabric · onelake · security · get-started-security
- fabric · onelake · onelake-shortcuts
- fabric · governance · governance-compliance-overview
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Fabric Workspaces, OneLake and Delivery Lifecycle
Memory hook: Capacity runs the work; workspace organizes it; OneLake stores it.
Must remember
Microsoft Fabric combines analytics experiences around OneLake and shared platform capabilities. A workspace contains items; capacity supplies supported execution resources. A domain organizes business ownership/governance across workspaces rather than acting as an automatic security boundary. Workspace assignment and settings influence where and how jobs run.
A Lakehouse combines files and managed table patterns, commonly Delta, with Spark and supported SQL access. A Warehouse targets relational T-SQL analytical development. Eventhouse/KQL databases target event/time-series analytical workloads. Choose from access pattern, transformation language, transaction needs, latency and operating model, not product name alone.
Configure Spark pools/session behavior, environments/libraries and workspace defaults consistently. OneLake settings affect supported access and integration. Apache Airflow workspace configuration enables its supported orchestration integration; credentials, dependencies and runtime settings still need control. Separate development and production assumptions.
Git integration versions supported item definitions; deployment pipelines promote supported content across stages. They are not automatically a complete backup of every underlying data file or secret. Use database projects for declarative SQL schema development where supported, review differences and manage environment-specific connections/parameters. A successful deployment does not prove target data or permissions are correct.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Spark-oriented Delta engineering | Lakehouse. |
| SQL-first relational analytics | Warehouse. |
| High-volume event exploration with KQL | Eventhouse. |
Traps
- A workspace is not the same as capacity.
- Git integration does not automatically version every byte of data.
02 · Fabric Security, Governance and Shared Data
Memory hook: Workspace access is broad; data rules need their own design.
Must remember
Separate workspace roles, item permissions and data-level controls. A user able to manage an item may have broader access than a consumer intended to see a subset. Apply supported row, column, object and folder/file controls at the layer and engine that actually serves the request. Test every access path, including shortcuts, SQL endpoints, Spark and reports.
OneLake security centralizes supported data-access policies, but feature scope, engine support and configured modes matter. SQL row-level predicates and column/object permissions solve different problems. Dynamic data masking changes displayed values for eligible queries; it is not encryption and not a substitute for restricting privileged/direct access.
Sensitivity labels communicate classification and apply supported protection/handling. Endorsement marks promoted or certified content under organizational governance. Neither replaces permission enforcement or validates business calculations. Audit logs record eligible actions for investigation; configure appropriate collection, retention and access.
Shortcuts reference data without an ordinary duplicate ingestion copy. Access to the shortcut and its target depends on the supported source/credential model. Validate trust boundaries, data residency and external connectivity. A shortcut can fail because its target, connection or permission changed even when its local definition remains intact.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Consumer sees only their business unit | Supported row-level rules with tested consumer access. |
| Hide a displayed sensitive value | Masking where appropriate, plus real authorization controls. |
| Share existing lake data without copying | A supported shortcut with validated target access. |
Traps
- Masked does not mean encrypted.
- One working query path does not prove every engine enforces identical permissions.
03 · Pipelines, Dataflows and Incremental Loads
Memory hook: Orchestrate dependencies; move data; make retries safe.
Must remember
Pipelines coordinate activities such as copy, notebooks and dataflows with dependencies, parameters and dynamic expressions. Dataflows Gen2 offer Power Query transformations; notebooks support code-oriented Spark and other supported execution. Choose the tool from complexity, language, scale and maintenance needs rather than assuming one tool fits every step.
Full loads replace or reload a complete dataset; incremental loads process changes since a tracked boundary. A watermark based on a timestamp/key must handle ties, updates, deletes and clock/source semantics. CDC captures changes from a supported source; mirroring replicates supported source changes into Fabric with product-specific limits. Neither guarantees a correct dimensional model automatically.
Use schedules when cadence is time-driven and event triggers when supported events should initiate work. Pass environment and partition parameters explicitly. Model retries, dependencies, concurrency and failure paths; a partially successful run must not duplicate or silently omit records. Stage data, validate it and publish atomically where the engine supports the required semantics.
Track source positions, row counts and business reconciliation totals. Treat late-arriving dimensions/facts deliberately, preserving unknown-member or correction logic. Protect credentials in managed connections and avoid writing secrets into parameters/logs. For Airflow-based orchestration, keep heavy data processing in the appropriate engine rather than passing large datasets through orchestration metadata.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Mostly visual cleanup and shaping | Dataflow Gen2. |
| Complex PySpark transformation | Notebook. |
| Coordinate copy and dependent transforms | Pipeline or appropriate Airflow workflow. |
Traps
- A timestamp watermark can miss records if its boundary logic is wrong.
- A successful copy is not proof that deletes and updates were handled correctly.
04 · Delta, PySpark, SQL and Dimensional Transformation
Memory hook: Preserve the grain before optimizing the engine.
Must remember
Raw, cleaned and curated layers often form a bronze/silver/gold pattern. These labels describe responsibilities, not mandatory product features. Delta tables add transactional metadata around supported lake files. Keep schema, partition layout and data quality explicit; simply placing Parquet files in a folder does not automatically create the intended managed Delta table.
Use PySpark for distributed code transformations, T-SQL for supported relational operations and KQL for event/time-oriented analytics. DataFrames are lazily evaluated until an action triggers work. Filter early, select needed columns and avoid collecting a large distributed dataset into a driver process. Joins can produce heavy shuffles or skew around hot keys.
Dimensional models define fact grain, dimensions and surrogate/business keys. Type 1 slowly changing dimensions overwrite the current attribute; Type 2 preserves history with new versions and effective ranges. Late facts need the dimension version valid at event time when historical correctness matters.
Deduplicate using a defensible key and ordering rule, not arbitrary row removal. Distinguish missing values from zero; standardize types/time zones; quarantine invalid records with enough context for repair. Group/aggregate only after deciding what detail can be discarded. Denormalization can improve consumption but may multiply data or complicate updates.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Need historical customer attribute at sale time | Type 2 dimension with correct effective-date join. |
| Current-only corrected spelling | Type 1 update may fit. |
| Driver out of memory | Inspect collect/to-local operations and distributed execution design. |
Traps
- Dropping duplicates without a deterministic rule can keep the wrong version.
- Aggregating early can destroy detail later analysis requires.
05 · Eventstreams, KQL and Streaming Windows
Memory hook: Time boundaries and late events decide the answer.
Must remember
Eventstream ingests, transforms and routes supported real-time events. Eventhouse provides KQL-oriented storage/query capabilities. Spark Structured Streaming offers code-based stateful processing. Select an engine from latency, state, transformation complexity, connectors and team skills.
Native Eventhouse tables ingest data into the engine; OneLake shortcuts reference external lake data. Query acceleration for supported OneLake shortcuts improves suitable query access through additional acceleration behavior and cost/refresh trade-offs. Choose after checking freshness, supported formats and performance needs; it is not identical to ordinary shortcut access.
Event time reflects when an event occurred; processing time reflects when it is handled. Tumbling windows do not overlap, hopping/sliding patterns may overlap, and session windows group activity by inactivity gaps where supported. Watermarks and late-data policies control how long state/results remain open to delayed events. Exact syntax and behavior differ between KQL, Eventstream and Spark.
KQL pipelines use operators such as where, project, extend, summarize and joins to shape event data. Apply selective filters early and choose appropriate time bins. Streaming checkpoints track progress/state for supported recovery, but external side effects still need idempotence. Monitor lag, dropped/late records and poison events rather than only whether the stream is running.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Visual event routing and supported transforms | Eventstream. |
| Fast event/time-series exploration | Eventhouse and KQL. |
| Complex stateful code transformations | Spark Structured Streaming with checkpoint/recovery design. |
Traps
- A running stream can be far behind real time.
- A window result can be incomplete if late data is discarded.
06 · Monitoring, Error Diagnosis and Performance
Memory hook: Locate the failing layer, then optimize the measured bottleneck.
Must remember
Monitor ingestion, transformation, semantic-model refresh and consumption. Track job state, duration, freshness, completeness, rejected records and capacity utilization. Alerts need owners and actionable thresholds. Capacity contention can slow several unrelated items at once, while a single malformed record may fail only one pipeline.
For pipeline/Dataflow failures, inspect the failing activity/step, parameters, credentials, gateway/connection, schema and source throttling. For notebooks, examine execution logs, library/environment versions, resource exhaustion and data paths. Eventstream/Eventhouse failures may involve source connections, schemas, transformation expressions or ingestion policies. T-SQL errors need actual statement/schema/permission evidence; shortcut errors need target/connection checks.
Optimize Lakehouse tables by addressing small files, partition design and supported compaction/layout options. V-Order targets suitable columnar read efficiency; it is not a universal remedy for every write-heavy workload. Cleanup such as VACUUM affects historical-file availability, so retention and readers matter. Do not delete underlying files manually to solve a table error.
Tune Spark partitions, join strategies, skew, caching and executor resources from observed stages. Tune warehouse queries by reducing unnecessary scans and expensive joins and checking concurrency. For streaming, examine throughput, batch/window/state behavior and backpressure. Pipelines benefit from appropriate parallelism, not unlimited simultaneous source requests.
Diagnosis drill: a daily dashboard is stale despite a green copy activity. Check transformation completion, target partitions, semantic-model refresh and report access in order. End-to-end correctness crosses more than one successful task.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Many jobs suddenly slower | Inspect shared capacity and concurrency. |
| One Spark stage dominates | Inspect skew, shuffle and partitioning before adding resources. |
| Lakehouse has many tiny files | Evaluate supported compaction and write-pattern changes. |
Traps
- More parallelism can overwhelm a source or capacity.
- A green upstream task does not prove a fresh downstream report.