certslothcertsloth
← DP-420 overview

Azure Cosmos DB AI Developer Associate / STUDY TOOLS

DP-420 quick review

This guide follows the refreshed Cosmos DB AI Developer scope, including vector/full-text retrieval, agent memory and Fabric integration. Older DP-420 outlines omit some of these areas.

Memory hook: Target the partition; measure the RU; protect the retrieval.

Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.

Scope/version: This guide follows the refreshed Cosmos DB AI Developer scope, including vector/full-text retrieval, agent memory and Fabric integration. Older DP-420 outlines omit some of these areas.

Must remember by domain

Domain Rapid revision
Model/throughput Account → database → container → item. Logical partition is determined by key value; physical partitions manage underlying distribution. Pick high-cardinality, balanced keys aligned with queries. Hierarchical keys support suitable tenant hierarchies; synthetic sharding spreads writes but can complicate reads. Manual/autoscale provisioned throughput and serverless suit different demand.
Consistency Strong prioritizes latest-read consistency; bounded staleness caps lag; session preserves session guarantees using tokens; consistent prefix prevents out-of-order observations; eventual converges. Stronger read guarantees can increase work/latency. Consistency is independent of the supported transaction boundary.
SDK operations Reuse a client, set preferred regions, target id+partition key for point reads and use continuation tokens for paging. ETag/If-Match rejects stale updates. Transactional batch is atomic within one logical partition; bulk optimizes throughput without global atomicity. TTL controls expiry, not exact job scheduling.
Change processing Processor distributes work with a lease container; pull mode gives direct control. Latest-version feed differs from all-versions-and-deletes history and prerequisites. Design at-least-once handling with idempotent side effects and post-success checkpoints. Copy-container jobs move supported data; AI-generated code still needs key/security review.
Security/recovery Entra data-plane RBAC differs from account management RBAC. Private endpoint/firewall permission is not data permission. Configure backups, restore rights, regional read/write behavior and conflict/failover policy. Periodic backup and continuous PITR differ; a replica cannot restore arbitrary past state.
Optimization/fleets Observe request charge, 429/retry-after, status/substatus, diagnostics and per-partition utilization. Point reads, partition filters, smaller projections and appropriate indexes reduce work. Composite indexes support selected query shapes. A hot partition may throttle despite account headroom. Fleet governance does not remove per-account/partition limits.
AI/analytics Full text = lexical matching; vectors = semantic similarity; hybrid combines them. Match dimension/type/metric/index to the embedding model. DiskANN/sharding trade recall, latency and scale. Agent state needs tenant/user boundaries, TTL and concurrency. Fabric mirroring and connectors support distinct analytics paths with their own freshness/security.

Diagnosis and traps

Locate the slow request → inspect charge/diagnostics → isolate client/network versus service time → inspect partition key/index/query → bound retries/concurrency → change capacity if evidence supports it. A high-cardinality random key can make required queries expensive. Vector similarity is not authorization; access filtering must precede evidence entering the model.

Last-pass self-check

1. Need atomic updates across arbitrary partition keys: transactional batch?

No. Its atomic boundary is one logical partition.

2. What should a stale ETag update do?

Fail its conditional check so the client can reload/reconcile instead of overwriting silently.

3. All item deletes needed: default latest-version feed enough?

No. Choose a supported mode/retention design that captures required deletes.

4. Why does RU scaling sometimes fail to solve throttling?

A hot partition, costly query or excessive retry concurrency can remain the bottleneck.

5. Why record embedding model/version?

Dimension or semantic-space changes can require regenerating vectors and indexes.

Sources

Every topic at a glance

Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.

01 · Resource model, consistency and partitioning

Memory hook: Good keys spread writes and target reads.

Must remember

  • An account contains databases and containers; items live within logical partitions determined by their partition-key values. Choose keys from access patterns, cardinality and write distribution, not only a field that looks unique.
  • Hierarchical partition keys use multiple key levels to support suitable tenant and scale patterns. Synthetic keys combine or distribute values deliberately; random sharding can reduce hotspots but complicate reads.
  • Embed related data when it is read/updated together and bounded; reference when independent growth or update patterns make duplication expensive. Model for the actual queries rather than copying a relational schema unchanged.
  • The five consistency choices are strong, bounded staleness, session, consistent prefix and eventual. Session consistency preserves supported session guarantees through tokens; consistency is not the same as transactional scope.
  • Request Units normalize operation cost. Provisioned manual/autoscale and serverless models suit different traffic patterns; evaluate account, database and container throughput settings and service constraints.
  • Use a long-lived SDK client where recommended, appropriate connection mode and region preferences. Capacity planning includes hot partitions, item size, indexing, query shape and concurrency.

Choose under exam pressure

Requirement Choice and reason
Most requests target one tenant’s records A tenant-aware key design, with growth/hot-tenant analysis and HPKs where suitable.
Occasional unpredictable traffic Evaluate serverless against provisioned/autoscale pricing and supported limits.

Traps

  • High total RU capacity cannot always rescue one hot logical partition.
  • A unique random partition key can make common multi-item queries expensive.

Practise this topic

02 · SDK operations, change feed and concurrency

Memory hook: Point read cheaply; retry safely; process changes twice safely.

Must remember

  • A point read uses item id and partition key and is usually cheaper than an equivalent query. Parameterize queries and use paging/continuation tokens for bounded result processing.
  • Create, replace, patch and delete have different update semantics. ETags with conditional requests implement optimistic concurrency so a stale writer does not silently overwrite a newer version.
  • Transactional batches operate within a supported logical partition boundary. Bulk execution improves throughput across operations but is not one global atomic transaction.
  • TTL expires eligible items according to container/item configuration; expiration is not a precise scheduling mechanism. Account for expired data in retention, change processing and recovery design.
  • Change feed supports reactive processing. The processor coordinates work through a lease container; pull mode gives the application more direct control. Latest-version versus all-versions-and-deletes modes have different prerequisites and event semantics.
  • Make consumers idempotent and checkpoint only after successful work. Copy-container jobs support suitable data movement; Agent Kit/AI-assisted scaffolding can accelerate development but generated partitioning, security and retry choices still require review.

Choose under exam pressure

Requirement Choice and reason
Fetch one known item Use id plus partition key for a point read.
Prevent lost updates Conditional writes using the current ETag and conflict handling.

Traps

  • Bulk writes do not make operations across arbitrary partitions atomic.
  • Change-feed processing must not assume a side effect will only ever be attempted once.

Practise this topic

03 · Secure, distribute and recover data

Memory hook: Identity plus network; replicas plus history.

Must remember

  • Prefer supported Entra identity and data-plane RBAC for applications, with narrowly scoped permissions. Control-plane roles and data access roles solve different problems; avoid distributing broad account keys.
  • Restrict network access using supported firewall/private endpoint controls and correct private DNS. Authentication still applies on a private path; test public-network settings independently.
  • Use encryption, auditing and supported data masking according to the API/feature. Dynamic masking changes what selected users see; it does not replace authorization or encryption against privileged access.
  • Multi-region reads improve locality; multi-region writes require conflict and consistency design. Configure preferred regions, failover priorities and supported partition/failover options, then test client behavior.
  • Periodic and continuous backup/PITR options have different capabilities and retention. Replication copies current state; recovery history is needed for accidental corruption or deletion.
  • Test restored data, identities, endpoints and application reconnection. Keep recovery permissions and retained backups usable while protecting production deletion paths.

Choose under exam pressure

Requirement Choice and reason
An application needs only one container Scoped data-plane permissions through a supported identity.
Recover a valid state before an accidental update A supported backup/PITR restore workflow.

Traps

  • A private endpoint does not grant database access.
  • Multi-region writes do not remove the need to understand conflicts and consistency.

Practise this topic

04 · RU efficiency, diagnostics and fleets

Memory hook: Measure per request and per partition.

Must remember

  • Inspect request charge, latency, status/substatus, query metrics and SDK diagnostics. Separate client/network delay from service processing and throttling.
  • Target partition keys, use point reads, reduce unnecessary fields and avoid unbounded scans. Index policies trade read efficiency against write/storage cost; composite indexes support particular query shapes.
  • Analyze index utilization and query plans before raising throughput. Large items, fan-out queries, excessive indexing and hot tenants can each drive RU consumption differently.
  • Azure Monitor metrics/alerts and diagnostic logs in Log Analytics connect service behavior to application SLOs. Monitor end-to-end latency and errors rather than only provisioned RU values.
  • Fleet capabilities manage supported groups of accounts, throughput policies and analytics across subscriptions. Fleet-level governance does not eliminate per-account/partition limits or the need to check feature support.
  • Load-test representative peak traffic, retries and regional behavior. Reuse SDK clients, respect retry-after guidance and cap concurrency to avoid amplifying a throttling incident.

Choose under exam pressure

Requirement Choice and reason
One tenant is throttled while the account has spare capacity Inspect partition distribution and tenant-specific load.
A query is expensive despite few returned rows Inspect scanned data, partition fan-out and index support.

Traps

  • Few returned records do not imply a cheap query.
  • More aggressive retries can worsen throttling and latency.

Practise this topic

05 · Vector retrieval, agent memory and Fabric

Memory hook: Retrieve with permissions; remember with boundaries.

Must remember

  • Full-text search matches lexical terms; vector search matches embedding similarity; hybrid search combines signals. Choose dimensions, data type, distance metric and index from the embedding model and workload.
  • DiskANN and supported sharded/vector-partitioned designs trade search performance, recall and scale. Evaluate filtered searches and index behavior using realistic tenant distributions.
  • RAG retrieves authorized, fresh evidence for generation. Chunking, embedding version, relevance and citations matter; a similar vector is not proof of a correct answer.
  • Store conversation state separately from durable semantic memory where appropriate. Use tenant/user keys, concurrency control, TTL/retention and deletion rules; do not let one user’s conversation become another user’s retrieval context.
  • Cosmos DB in Fabric and Azure Cosmos DB serve different integration/operational choices. Mirroring exposes supported operational data for Fabric analytics; verify lag, schema handling and security.
  • Fabric T-SQL/Spark can analyze mirrored JSON data, while the Cosmos DB Spark connector supports appropriate reads/writes. Analytics access must preserve data governance and avoid overwhelming the operational workload.

Choose under exam pressure

Requirement Choice and reason
Need exact product codes and semantic descriptions Evaluate hybrid lexical/vector retrieval with filters.
Analyze operations without repeated application queries Supported Fabric mirroring with governance and freshness checks.

Traps

  • An embedding model change may require re-embedding and index compatibility review.
  • Vector-only memory can lose exact identifiers, chronology and authorization context.

Practise this topic

Search across every published topic.