certslothcertsloth
← AI-500 overview

Multi-Agent AI Solutions Expert / STUDY TOOLS

AI-500 quick review

Memory hook: One agent boundary, one accountable action.

Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.

Must remember by domain

Domain Rapid revision
Architecture Decompose the goal into deterministic steps, bounded agent decisions and tools. Specify each persona, autonomy, input/output, owner, permissions and exit. Sequential suits dependencies; parallel suits independent tasks; hub/orchestrator patterns centralize coordination; peer-to-peer needs explicit conflict and handoff rules.
Platform Match models/compute to quality, latency, concurrency, state lifetime and cost. Design per-agent identity, tenant boundaries, durable state, trace correlation and recovery before coding. Human-AI experience includes understandable approval, correction, override and escalation.
Prompt and memory Version instructions/examples/dynamic context. Separate session state, shared team state and long-term semantic memory. Retrieval/summarization/compaction must preserve exact identities and commitments. Sliding-window loss, summary drift and vector-only recall are different failure modes.
Tools and protocols Function calling proposes arguments; trusted application code validates/authorizes execution and validates results. MCP exposes supported tools/resources; A2A connects agents. Functions, Logic Apps and API Management can expose controlled operations. Use schemas, scoped credentials, timeouts, idempotency and explicit errors.
Orchestration Agent Framework, LangGraph/LangChain and Transformers have distinct responsibilities. Central middleware handles authorization/logging/errors. Cap agent spawning, concurrency, retries, tokens and time; propagate cancellation. Prompt cache, exact-response cache and semantic cache reuse different data and require access/freshness boundaries.
Evaluation and operations Score memory/retrieval quality, tool choice/arguments, trajectory and final outcome separately. Calibrate human/model judgments with held-out/adversarial examples. Trace latency and cost across all child agents; a correct answer obtained through unauthorized actions fails the task.
Security and delivery Prefer workload identity and supported OAuth/on-behalf-of flows; scope Key Vault access. Apply guardrails at input, retrieval, tool request, tool response and output. Red-team indirect injection and cross-tenant leakage. Promote through tested environments with canary/blue-green rollback of code, prompts, tools and compatible state.

Failure sequence and traps

Find the first divergent trace → inspect authorized input/state → check retrieval → check tool schema/result → check handoff/termination → replay a safe representative test. More agents can increase errors and cost. A natural-language instruction to ask approval is weaker than a workflow state that prevents execution until approval exists. Capture observable decisions and tool events; do not assume access to private internal model reasoning.

Last-pass self-check

1. A summary drops an account ID: which failure?

Context compaction/summary drift; preserve exact structured identifiers outside lossy prose.

2. Why validate tool output as well as input?

It can be malformed, false, stale or contain indirect instructions.

3. Which cache is safe across tenants by default?

None containing private results; design keys, authorization and retention explicitly.

4. Can an LLM judge be the only production gate?

No. Calibrate it and combine task, policy and human review where required.

5. What must stop when the parent task is canceled?

Dependent agents/tools and retries, with side effects reconciled safely.

Sources

Every topic at a glance

Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.

01 · Architect bounded multi-agent systems

Memory hook: One responsibility, one boundary, one exit.

Must remember

  • Decompose goals into deterministic workflow steps, agent decisions and tool actions. Use an agent only where flexible reasoning/interaction adds value; fixed business rules can stay in ordinary code.
  • Define each agent’s role, allowed tools, data scope, autonomy, outputs and termination criteria. Human-AI experience design includes understandable approvals, correction, escalation and user control.
  • Sequential, parallel, hub-and-spoke, peer-to-peer and orchestrator/subagent patterns have different failure and coordination costs. Match dependencies and ownership rather than adding agents for appearance.
  • Select model families and compute from task quality, latency, concurrency, state and cost. Foundry, Azure Functions, Container Apps/other supported compute, data stores and integration components must fit the execution lifetime.
  • Design session, shared-team and long-term memory separately with tenant boundaries and retention. Zero Trust requires per-agent identities and constrained lateral access.
  • Specify trace correlation, replayable observable events, health metrics and quality checks during design. Use reproducible dev containers, dependencies, CLI/editor tooling and reviewed AI instruction files for the development workflow.

Choose under exam pressure

Requirement Choice and reason
An approval must always precede a payment A deterministic authorization gate around the agent’s proposed action.
Several independent analyses feed one decision Bounded parallel workers with a coordinator that validates and reconciles results.

Traps

  • A multi-agent design can increase latency, cost and failure modes.
  • A persona written in a prompt is not an enforceable permission boundary.

Practise this topic

02 · Prompts, context and durable memory

Memory hook: Context is selected evidence, not a transcript dump.

Must remember

  • Use clear role/task instructions, examples, output schemas, dynamic context and defensive boundaries. Version prompts and evaluate changes across representative tasks and adversarial inputs.
  • Short-term context supports the current session; long-term stores preserve selected facts or semantic memories. Apply authorization, provenance, expiry/deletion and tenant isolation to every read/write.
  • Context accumulation eventually exceeds a window. Retrieval, summarization and compaction can help, but preserve exact identifiers, commitments and authoritative source references needed for continuity.
  • Sliding-window amnesia loses older facts; summary drift distorts meaning; vector-only recall can miss exact entities or chronology. Combine structured state with semantic retrieval where appropriate.
  • Multi-agent RAG requires suitable chunking, embedding versions, metadata/security filters and retrieval precision. Search/MCP-accessible sources must preserve the requesting identity’s access rights.
  • Fine-tuning is a separate controlled lifecycle with curated data, evaluation and retraining frequency. Use it for measured behavior needs rather than as a replacement for current factual retrieval.

Choose under exam pressure

Requirement Choice and reason
An agent repeatedly forgets a customer identifier Store exact structured session state instead of relying only on semantic memory.
Agents need current policy evidence Authorized retrieval with source versions and relevance evaluation.

Traps

  • Summarization can discard a constraint that matters later.
  • Shared vector memory can leak data unless access is enforced before retrieval.

Practise this topic

03 · Tools, MCP and orchestration

Memory hook: Validate every tool; bound every loop.

Must remember

  • Function calling proposes structured arguments; application code authorizes, validates and executes the operation. Check both input and tool-result schemas and reject unsafe or malformed requests.
  • MCP servers/clients expose supported tools and resources; A2A connects agent systems. Authentication, user consent, tenant isolation and network policy remain application responsibilities.
  • Functions, Logic Apps and API Management can expose controlled capabilities. Use allowlisted operations, scoped identities, timeouts, idempotency keys and explicit error contracts.
  • Agent Framework, LangGraph/LangChain and other frameworks express orchestration differently; select by required state, branching, integration and operational support. Hugging Face Transformers supplies model capabilities rather than automatically solving workflow governance.
  • Implement approval, override and escalation as explicit states. Cap spawned agents, concurrency, token use, retries and elapsed time; cancel downstream work when the parent workflow stops.
  • Prompt caching, response caching and semantic caching cache different things. Include model/prompt version, tenant/access context and relevant freshness in cache design; do not reuse a private answer across users.

Choose under exam pressure

Requirement Choice and reason
A tool retries after a network timeout Use a stable idempotency key and inspect whether the action already succeeded.
A semantically similar query comes from another tenant Re-evaluate authorization; do not return a shared cached private response.

Traps

  • MCP discovery does not authorize unrestricted execution.
  • Parallelism can exceed model/API quotas and make the whole workflow slower.

Practise this topic

04 · Evaluate, monitor and control cost

Memory hook: Outcome, trajectory, quality, cost.

Must remember

  • Evaluate memory retrieval, knowledge relevance, prompt behavior, tool choice/arguments and final task completion separately. A correct final answer can hide an unauthorized or wasteful trajectory.
  • Use human review, built-in/custom metrics and calibrated LLM judges with golden and adversarial datasets. Validate synthetic test data and avoid judging only examples used during development.
  • Trace observable model, agent and tool spans with correlation IDs. Foundry tracing and structured logs should show execution, timing, tokens, errors and approved replay context without indiscriminate sensitive-data capture.
  • Monitor agent health, coordination failures, drift, quality regressions and service availability. Define SLOs for successful task completion and latency, with runbooks for recurring failure patterns.
  • Optimize bottlenecks through model selection, bounded parallelism, retrieval, caching and prompt length. Respect rate limits; track tokens, tool calls, quotas, allocations and chargeback at useful ownership boundaries.
  • Use feedback loops and controlled A/B/canary comparisons. An automated improvement loop must still pass quality, safety and cost gates before changing production.

Choose under exam pressure

Requirement Choice and reason
A workflow succeeds but costs ten times more Inspect repeated tool/model calls, context growth and runaway coordination.
Quality drops after a prompt release Compare versioned evaluation results and traces, then roll back if gates fail.

Traps

  • Low latency is not success if the agent silently skips required work.
  • An LLM judge’s score is not a substitute for calibrated evidence.

Practise this topic

05 · Zero Trust, guardrails and release

Memory hook: Policy outside prompts; release behind gates.

Must remember

  • Use per-agent/workload identity, scoped RBAC and network boundaries. OAuth/on-behalf-of flows preserve supported user context; API keys require secure storage and rotation when unavoidable.
  • Key Vault stores supported secrets, keys and certificates with access controls and lifecycle management. Do not copy a broad secret into every agent environment.
  • Apply guardrails at user input, retrieval, tool request, tool response and final output. Domain-specific constraints such as transaction limits require deterministic validation, not only content filtering.
  • Red-team the system early, including indirect prompt injection, data exfiltration, cross-tenant access and unauthorized actions. Foundry red-team tooling can support tests; inspect coverage and validate mitigations.
  • Development-Test-Acceptance-Production, canary and blue/green strategies control release exposure. Version infrastructure, agents, prompts, tools and data contracts; test rollback with compatible state.
  • CI/CD needs unit, integration, regression and evaluation gates plus scoped release identities. Human approval must describe the exact proposed consequential action, not a blanket permission for future unknown operations.

Choose under exam pressure

Requirement Choice and reason
An external document asks the agent to reveal keys Treat it as untrusted content and enforce tool/data policy independently.
Release changes both tools and prompts Promote a tested versioned bundle with compatibility and rollback checks.

Traps

  • A system prompt is not an access-control list.
  • Input screening alone misses malicious content returned by tools or retrieval.

Practise this topic

Search across every published topic.