Memory hook: Image → runtime → data → events → trace.
Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.
Scope/version: AI-200 is the current AI Cloud Developer path. AZ-204 retired in July 2026; its older objectives are not an exact match.
Must remember by domain
| Domain | Rapid revision |
|---|---|
| Containers | ACR stores images; ACR Tasks build/automate them. App Service hosts supported web containers; Container Apps provides revisions, ingress and KEDA scaling; AKS exposes Kubernetes orchestration. Pin image digests, configure environment/secrets and keep registry-pull identity separate from application identity. |
| Kubernetes operation | Deployment manages replicas/rollouts; Service selects matching ready endpoints. Readiness controls traffic; liveness restarts failed containers; startup probes protect slow initialization. Requests affect scheduling; limits constrain resources. Pending means investigate scheduler events; CrashLoopBackOff means investigate process/configuration/dependencies. |
| Cosmos DB | Use a reused SDK client, parameterized queries, partition-key targeting and suitable indexing/consistency. A point read uses both id and partition key. RU charge and per-partition hotspots explain throttling better than account averages. Vector dimensions/metric must match embeddings. Change feed processors use leases and require idempotent handlers. |
| PostgreSQL | Model types, keys and indexes for actual queries. pgvector supports semantic retrieval with exact/approximate choices; filtering and ANN settings affect recall. Inspect plans and memory/IO before resizing. Pool connections and bound concurrency; too many new connections can dominate latency. |
| Managed Redis | Cache-aside reads cache, fetches on miss and populates with expiry. Coordinate invalidation after writes; TTL bounds age without guaranteeing immediate freshness. Vector indexes enable similarity retrieval. Tenant, model and permission context belong in cache design. |
| Messaging and Functions | Service Bus queues distribute work; topics/subscriptions fan out brokered messages. Peek-lock processing completes after success; expired/abandoned locks can cause redelivery. Dead-lettered messages need investigation and explicit replay. Event Grid distributes events with filters/retries. Functions has one trigger plus optional input/output bindings. |
| Security and observability | Key Vault holds secrets/keys/certificates; App Configuration holds application settings and feature flags. Use managed identity and scoped permissions. OpenTelemetry correlates traces across services; KQL filters and aggregates logs. Configuration access and secret access are separate grants. |
Diagnose in order
Ingress/DNS → Service selectors → ready pod/revision → dependency route → authentication → data/query behavior. For queues, check arrival rate, processing time, retries and downstream capacity before raising replica limits. Event-driven delivery can repeat: use idempotency and bounded retries, never assume every trigger means a unique business operation.
Last-pass self-check
1. A pod runs but gets no traffic: first checks?
Readiness, matching labels/selectors, endpoints and the listening port.
2. Which service suits durable work with dead-letter handling?
Service Bus; Event Grid primarily distributes event notifications.
3. What does a 429 from Cosmos suggest?
Inspect throttling/retry-after, request cost, partition hotspots and concurrency.
4. Does a cache TTL make writes immediately visible?
No. Use a deliberate invalidation/update policy.
5. Why follow one trace ID across services?
To separate gateway, runtime, database and model/tool latency or failures.
Sources
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · App Service and Container Platforms
Memory hook: Separate image storage, execution, application configuration and the hosting plan.
Must remember
- ACR stores container images/artifacts. Authenticate pushes/pulls with appropriate identities and roles; prefer immutable image digests for a known release. Registry access does not grant the running application permission to its database.
- Container Instances runs container groups without managing a cluster. Container Apps provides managed application environments, revisions, ingress and event-driven scaling capabilities. AKS gives Kubernetes orchestration with greater platform control and responsibility; it is not the default answer for every container.
- Size container CPU/memory and configure health/startup behaviour. Container Apps scaling rules and minimum replicas affect availability, cold starts and cost. Separate revision traffic from image publishing; pushing an image alone is not necessarily deployment.
- An App Service plan determines shared compute capacity, Region and pricing tier; apps run within it. Scale up changes plan capability; scale out changes instance count. Apps sharing a plan can compete for its resources.
- Deployment slots support staged releases and swaps on eligible tiers. Mark environment-specific settings as slot settings where needed. Warm up and validate the target before swapping; external database changes require their own compatibility plan.
- Configure custom-domain ownership, DNS records and TLS bindings separately. VNet integration primarily handles supported outbound access; private endpoints support private inbound access. Access restrictions, DNS and the destination's permissions still matter.
- Backup support depends on plan/features and configuration; define a restore test. Managed identity and Key Vault references reduce stored credentials. Inspect app logs and dependency/network failures before simply increasing plan size.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Run a small container without managing nodes | Container Instances or Container Apps, according to application/scaling needs. |
| Test a web release before moving users | App Service deployment slot. |
| App needs private database access | Supported VNet integration plus routes, DNS and authorisation. |
Traps
- A registry is not a container runtime.
- VNet integration is not equivalent to private inbound access.
- A slot swap does not reverse database writes.
02 · Microsoft Entra ID and Azure RBAC
Memory hook: Entra identifies the principal; Azure RBAC authorises an action at a scope.
Must remember
- A Microsoft Entra tenant is an identity directory. An Azure subscription is a billing/resource-management boundary associated with a tenant. Management groups organise subscriptions; resource groups organise resources. Do not treat tenant and subscription as synonyms.
- Manage users, groups, properties, assigned licences and guest access deliberately. Security groups organise access; dynamic membership uses rules when licensing/features permit. B2B guests retain an external identity relationship. Self-service password reset needs appropriate eligibility, authentication methods and configuration.
- An Azure role assignment is principal + role definition + scope. Scope can be management group, subscription, resource group or resource, with inheritance. Inspect effective assignments rather than only the nearest resource. Entra directory roles and Azure resource roles are different permission systems.
- Owner can manage resources and access; Contributor manages resources but does not normally grant Azure roles; Reader reads management information. Data-plane roles, such as Storage Blob Data Reader, authorise data operations separately. Management-plane access is not always data access.
- System-assigned managed identity follows one resource's lifecycle; user-assigned identity is an independent reusable resource. Both avoid embedded secrets for supported authentication. Grant the identity's service principal only the required target roles.
- PIM supports time-bound/eligible privileged access; Conditional Access evaluates sign-in conditions and grant controls with appropriate licensing. MFA strengthens authentication; it does not create a missing resource role assignment. Keep a monitored emergency-access design.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| App needs storage access without stored credentials | Managed identity plus an appropriate data role. |
| User can manage a storage account but cannot read blobs | Check data-plane permissions. |
| Temporary privileged operations | Eligible/time-bound access using PIM where available. |
Traps
- Entra administrator is not automatically Owner of every Azure subscription.
- Contributor and Owner differ in access-management privileges.
- A role at a parent scope can remain effective after a narrower assignment is removed.
03 · Azure Monitor, Logs and Alerts
Memory hook: Metrics quantify, logs explain, traces connect and alerts start a response.
Must remember
- Azure Monitor combines metrics and logs; Log Analytics workspaces hold queryable log data. Activity Log records management-plane events; resource/application logs require relevant collection configuration. Diagnostic settings route supported categories to selected destinations.
- Azure Monitor Agent uses data collection rules for supported guest telemetry. VM, Storage and Network Insights provide focused views; Application Insights adds application performance and tracing with suitable instrumentation. Guest memory/disk metrics are not automatically identical to platform metrics.
- KQL pipelines transform tables:
wherefilters,projectselects columns,summarizeaggregates,bin()groups time intervals, and joins combine data. Example:Heartbeat | summarize LastSeen=max(TimeGenerated) by Computerfinds each computer's latest recorded heartbeat; absence can mean collection failure, not only host failure. - Alert rules define signal, scope, evaluation and condition. Action groups define notifications/actions. Alert processing rules modify processing such as suppression under selected conditions; they do not change the source telemetry. Use dynamic thresholds where appropriate and test missing-data behaviour.
- Network Watcher and Connection Monitor help inspect path and connectivity. Logs, effective routes/rules and an application test answer different questions. Narrow time windows and correlation IDs reduce noise; retention and ingestion volume affect cost.
- Monitor the collection pipeline itself. Permissions, network access, workspace configuration and data collection rules can break visibility. Avoid logging secrets and unnecessary personal data, and define retention according to operational/evidence requirements.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Who changed an Azure resource? | Activity Log and relevant audit evidence. |
| Notify an operations group on a metric breach | Alert rule linked to an action group. |
| Investigate recurring connection failures | Connection Monitor plus route, rule and application evidence. |
Traps
- No logs can mean no collection.
- Action groups do not define the alert threshold.
- Average response time can conceal severe tail latency.
04 · RAG, Agents and Production Integration
Memory hook: Authorise evidence before retrieval, and authorise actions before tool execution.
Must remember
- Build ingestion with source permissions, stable IDs, text/layout extraction, chunking, embeddings and index updates/deletes. Azure AI Search supports lexical, vector and hybrid retrieval with supported semantic ranking/enrichment capabilities. Vector similarity, ranking and grounded answer quality are different measurements.
- Enforce tenant/document permissions using trusted identity context before evidence enters the model. Preserve source references for citations and freshness. Tune chunk size, overlap, filters and reranking on held-out questions; more retrieved text is not always better.
- Agents need explicit roles/goals, conversation state, memory, typed tool schemas and limits. Integrate APIs, search, knowledge stores, custom functions and content-analysis tools through supported interfaces. A tool result is untrusted input and may itself contain injection.
- Multi-agent systems require ownership of tasks/shared state, clear handoffs, failure propagation and end-to-end evaluation. Supervisor/delegation patterns add cost and latency. Prefer a deterministic workflow where known branches are sufficient; use rules for exact constraints rather than relying on model obedience.
- Record project/deployment configuration, prompts, index versions and tool contracts in CI/CD. Test with managed identities/private networking and the actual application roles. Treat model, prompt and retrieval changes as releases with canary/rollback and compatibility checks.
- Monitor tokens, quotas, time to first useful output, completion latency, tool failures and cost per successful task. Bound retry/reflection loops, validate structured output, cache with identity/freshness context and evaluate smaller-model routing where quality permits.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Current internal knowledge with citations | RAG with ACL-aware retrieval and support checks. |
| Several tools need controlled sequencing | A workflow or bounded agent orchestration. |
| A model requests an irreversible action | Validate policy and required approval before execution. |
Traps
- Reflection can repeat an error rather than correct it.
- A cache without tenant context can leak data.
- Changing an embedding model can require index migration.
05 · AKS Manifests and Container Scaling
Memory hook: Desired replicas, healthy pods, service discovery and ingress are separate checks.
Must remember
- A Kubernetes Deployment declares pod templates and rollout state; a Service supplies stable discovery/traffic selection; ConfigMaps and Secrets supply configuration with different intent. A Secret object is not a substitute for encryption, least privilege or external secret lifecycle.
- Match labels/selectors and container ports carefully. Readiness controls eligibility for traffic; liveness can restart a stuck container; startup probes protect slow startup from premature liveness failure. Incorrect probes can create a restart loop.
- Requests guide scheduling; limits constrain supported resource use. Pending pods may lack capacity or satisfy no node/affinity constraints. CrashLoopBackOff needs container logs/events and configuration/dependency checks, not only more replicas.
- ACR Tasks automate supported image builds/operations. Use immutable image versions/digests and scoped image-pull identity. Application identity is separate from registry pull permission; use supported workload identity for Azure API access.
- Container Apps revisions and KEDA-based rules support event-driven scaling. Define minimum/maximum replicas and suitable signals/credentials. Scaling on queue depth must account for processing time, retries and downstream capacity.
- Diagnose from ingress to service selectors/endpoints to ready pods to dependency DNS/network/authentication. Use logs/events and distributed traces. A successful image pull does not prove the application's listener or credentials work.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Pod runs but receives no traffic | Check readiness, labels/selectors and service endpoints. |
| Pod stays Pending | Inspect events, resources and scheduling constraints. |
| Queue load should drive replicas | A suitable KEDA/event scaling rule with bounded capacity. |
Traps
- Liveness and readiness are not interchangeable.
- A Kubernetes Secret can still be exposed through logs or broad RBAC.
- More replicas can overload the database.
06 · Cosmos DB, PostgreSQL Vectors and Managed Redis
Memory hook: Choose keys and indexes for the query, then make cache freshness explicit.
Must remember
- Cosmos DB for NoSQL SDK operations need endpoint/authentication, database/container and partition-key context. Point reads by ID and partition key differ from cross-partition queries. Request Units reflect work; examine query metrics, indexing policy and chosen consistency before adding throughput.
- Store embeddings with compatible dimensions and supported vector indexing/query configuration. Filter by trusted tenant/metadata constraints. Vector similarity is not access control or proof of answer correctness. A model/dimension change can require index/data migration.
- A change feed processor tracks supported changes through leases/checkpoints and distributed workers. Design idempotent handlers and verify the feed mode's treatment of updates/deletes; do not assume every mode captures every historical operation.
- PostgreSQL needs sensible tables/types, keys and indexes plus connection pooling. pgvector supports vector similarity operations and index choices with recall/performance trade-offs. Metadata predicates, candidate count, index parameters, memory and compute affect latency and quality.
- Bound connection counts; reuse pools safely and set timeouts. A larger database can still be bottlenecked by poor queries or too many short-lived connections. Inspect query plans and resource metrics before scaling.
- Azure Managed Redis supplies cache operations and supported vector search capabilities. Set TTL, eviction and invalidation deliberately. Cache-aside loads on a miss; write-through updates a cache during writes. Protect against stampedes and include identity/context in sensitive result cache keys.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Known Cosmos item ID and partition key | Prefer a point read where it meets the requirement. |
| Embedding search is slow | Inspect vector index, filters, candidate settings and resource limits. |
| Frequently requested data changes occasionally | Cache with a defined invalidation/TTL strategy. |
Traps
- A cache is not automatically the source of truth.
- Cosmos change-feed modes have different semantics.
- Vector dimensions must match the chosen embedding/index configuration.
07 · Service Bus, Event Grid and Functions Integration
Memory hook: Messages need settlement, events need routing, functions need bounded retry-safe execution.
Must remember
- Service Bus queues deliver work to competing consumers; topics/subscriptions fan out with subscription filters. Peek-lock processing needs completion, abandonment/dead-lettering or lock renewal as appropriate. Duplicate detection and sessions solve specific deduplication/ordering needs; still make side effects idempotent.
- Inspect dead-letter reason/description, fix the cause and replay deliberately. Poison messages must not cycle forever. A lock timeout can cause redelivery after the handler has already changed an external system.
- Event Grid routes supported/custom events with filtering, retry and dead-letter configuration. Validate event schemas and subscription handshakes; do not treat an event notification as a durable authoritative database record by itself.
- Functions triggers start execution; bindings connect supported input/output services. Build HTTP APIs with authentication, validation and useful error responses. Hosting plan affects scale, execution, networking and cost; deployment settings and runtime versions must match the code.
- Key Vault holds secrets/keys/certificates with controlled retrieval/rotation. App Configuration holds application settings and feature flags; it is not simply another password store. Use managed identities and scoped roles rather than embedded keys.
- OpenTelemetry propagates trace context across HTTP, queues and functions when instrumented correctly. Correlate logs with spans; use KQL to inspect dependency failures and latency. Redact secrets and payloads, monitor retries/locks/lag and distinguish invalid input from transient service failures.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Durable business work with competing workers | Service Bus queue. |
| Route notifications to interested handlers | Event Grid. |
| One slow operation risks message-lock expiry | Renew appropriately or redesign bounded processing with idempotency. |
Traps
- Completing a message before durable success can lose work.
- A binding simplifies code but does not remove permissions or failure semantics.
- Retrying every exception can duplicate actions and hide permanent errors.