certslothcertsloth
← AI-103 overview

Azure AI Apps and Agents Developer Associate / STUDY TOOLS

AI-103 quick review

AI-103 is the current Apps and Agents Developer path. AI-102 retired in June 2026; this guide follows the newer Foundry and agent objectives.

Memory hook: Ground the answer; constrain the action.

Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.

Scope/version: AI-103 is the current Apps and Agents Developer path. AI-102 retired in June 2026; this guide follows the newer Foundry and agent objectives.

Must remember by domain

Domain Rapid revision
Plan and manage Choose large/small/multimodal models and Foundry Tools by required capability, region, throughput and cost. Provision identity, network access, deployments and connected resources through repeatable configuration. Quotas, rate limits and provisioned capacity are distinct constraints.
Governance and operation Evaluate quality, safety, subgroup behavior and authorized task completion. Version model, prompt, index, tools and evaluation data. Trace retrieval/model/tool spans; retain approvals and provenance while redacting sensitive content.
Generative applications RAG sequence: ingest/extract → chunk → embed/index → apply trusted access filters → retrieve/rerank → generate with evidence → validate/cite. Hybrid retrieval combines lexical and vector signals; semantic ranking reranks suitable candidates. Missing evidence and misuse of good evidence need different fixes.
Agents Define role, tools, typed arguments, state, memory, handoffs and termination. Function calling proposes an operation; application code authorizes it. Sequential/parallel/orchestrator patterns trade coordination, latency and cost. Rules should enforce exact constraints outside prompts.
Vision generation Text/reference media can drive image/video generation where supported. Inpainting uses a mask for selected edits. Check model-specific controls and watermark/brand policy; an image-input endpoint is not automatically an image-output endpoint.
Visual understanding Ground captions and visual Q&A in the actual image/video segment. Alt text should serve accessibility; location, bounding regions and timestamps preserve verification. Embedded text can carry indirect prompt injection.
Text and speech Extract entities/topics, classify tone/sentiment, summarize, translate and validate structured JSON. Speech recognition, synthesis and translation have different directions; custom speech addresses supported domain/acoustic requirements. Test real accents, vocabulary and noise.
Information extraction OCR reads text; layout preserves structure; Content Understanding analyzers extract configured fields/representations. Built-in/custom search skills enrich ingestion. Select single-task/pro-mode capabilities deliberately and await long-running results before consuming them.

Diagnostic order and traps

For a bad answer: source freshness/permissions → extraction → chunking/index → retrieval relevance → prompt/context → model output. For a slow agent: trace dependency spans, retries, token use and loop termination before changing the model. Reflection is another model pass, not proof of correctness. Never place untrusted document instructions above system policy or share cached private answers across identities.

Last-pass self-check

1. Retrieval found no relevant evidence: fine-tune first?

No. Repair ingestion, permissions, retrieval or coverage before tuning answer style.

2. Can a valid citation still support a wrong answer?

Yes. Check that the cited passage actually entails the claim.

3. Why retain model and index versions?

Either can change behavior independently of application code.

4. How should consequential actions be controlled?

Validate schema, authorize scope and obtain required approval before executing the tool.

5. Why evaluate multimodal safety separately?

Images, audio and video introduce different unsafe content, embedded instructions and accessibility failures.

Sources

Every topic at a glance

Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.

01 · AI Models, Workloads and Responsible Use

Memory hook: Identify the input, output and consequence before selecting a model.

Must remember

  • Machine learning predicts patterns from data; generative models create content; agents combine models with state and tools to pursue tasks. Supervised classification predicts categories, regression numbers, clustering finds groups, and reinforcement learning improves behaviour through rewards.
  • Foundation models support broad tasks after pretraining. Large/small models trade capability, latency, footprint and cost. Multimodal models accept or produce supported combinations of text, images, audio or video. Model capability and deployment availability vary by version and Region.
  • Tokens are model-specific units; context/output limits constrain requests. Temperature and related sampling settings influence variability, not truth. Embeddings represent content for similarity retrieval; they are not encryption or final answers.
  • Choose text analysis for entities, key phrases, sentiment or summaries; speech recognition for audio-to-text; synthesis for text-to-audio; vision for visual interpretation; generation for new media; extraction for structured information from source content. A fixed exact calculation may be safer as ordinary code.
  • Responsible AI considers fairness, reliability/safety, privacy/security, inclusion, transparency and accountability. Evaluate representative groups and failure costs. Explain intended use and limitations, offer accessible experiences and assign a responsible owner for correction/escalation.
  • Hallucinations can be fluent and wrong. Ground responses in suitable evidence, validate structured outputs and use human review where consequence requires it. Filters reduce selected risks but cannot guarantee correctness, fairness or legal suitability.

Choose under exam pressure

Requirement Choice and reason
Classify an incoming request A suitable classifier/model with measured category performance.
Answer using current company documents Retrieval-grounded generation with access controls.
Perform an exact regulated calculation Deterministic validated logic when prediction is unnecessary.

Traps

  • Model confidence is not independent evidence.
  • A larger model is not always the best operational choice.
  • Removing a sensitive column does not remove every proxy for that attribute.

Practise this topic

02 · Foundry Projects, Prompts and Lightweight Clients

Memory hook: Authenticate, select the deployment, send a bounded request, inspect the result and handle failure.

Must remember

  • A Foundry project organises supported AI application resources/connections. Choose a model and deployment option using capability, Region, quota, throughput, latency and cost. A catalogue model name and your deployment identifier are not always interchangeable.
  • In the portal, deploy an available model, test representative prompts and inspect output/usage. A successful playground example does not validate application authentication, networking or error handling. Record the model/deployment and configuration used.
  • System instructions define application behaviour; user input provides the request; retrieved text/tool output is untrusted data. Use clear tasks, context, constraints and output schemas. Zero-shot uses no examples; few-shot includes demonstrations. Keep prompts versioned with evaluation cases.
  • A lightweight client needs the supported SDK, endpoint/project configuration, a credential and deployment/model selection. Prefer Entra/default-credential patterns for supported keyless access, with the required roles. Create a client, submit messages/input, inspect text/structured results and usage, and handle authentication, rate-limit and timeout errors.
  • Streaming returns incremental output; it requires correct accumulation, cancellation and partial-failure handling. Never put keys in source code. SDK shapes evolve, so use the current language quickstart for exact imports and methods rather than memorising an old preview signature.
  • A single agent adds a goal/instructions, tools and conversation state. Test it in the portal, then integrate the supported agent client lifecycle: create/reference the agent, submit a user turn, process required tool actions under policy and collect the final result. Bound loops and verify real tool outcomes.

Choose under exam pressure

Requirement Choice and reason
Prototype model suitability Portal deployment/playground with representative tests.
Production app needs credentials without a stored key Supported Entra/managed-identity authentication.
Model requests a tool action Validate and authorise the action before execution.

Traps

  • Deployment ID and base model name may differ.
  • A successful portal call does not prove the app identity has permission.
  • Agent text saying “done” is not proof an external action succeeded.

Practise this topic

03 · Text, Speech, Vision and Content Understanding

Memory hook: Interpret existing evidence, generate new content and extract structured fields as different tasks.

Must remember

  • Text analysis includes entity/key-phrase extraction, sentiment, summarisation and sensitive-content detection. Use supported Foundry Tools or model prompting according to accuracy, format and control requirements. Validate returned JSON/types before using results downstream.
  • Azure Speech supports speech recognition and synthesis; translation and multimodal audio models address related but different tasks. Match language, audio format, streaming/batch mode and latency. A spoken prompt can be transcribed before a text model or sent to a compatible multimodal model.
  • Vision-capable models interpret image inputs for captions, questions or visual evidence. Image-generation models create new visual outputs from prompts/references. Supported models may offer editing/masks; accepting an image does not mean a model can generate one.
  • Content Understanding uses configured analyzers to extract information from supported documents, images, audio and video. Define desired fields/schema, submit source content and consume the resulting structured output/evidence. Layout, timestamps and provenance can matter as much as the extracted value.
  • Long-running analysis may return an operation identifier; poll/await completion according to the SDK rather than treating acceptance as a final result. Validate confidence/evidence and route uncertain high-impact fields for review. An extracted invoice total still needs a business validation rule.
  • Protect input data, apply content/safety controls and preserve consent/licensing. Embedded text in images or documents can contain prompt injection. Accessibility captions should describe useful visible information without inventing details. Measure performance on the actual languages, document layouts and recording conditions.

Choose under exam pressure

Requirement Choice and reason
Extract invoice fields into a schema A suitable Content Understanding analyzer.
Generate spoken output Speech synthesis.
Answer a question about an existing image A vision-capable multimodal model with grounded evaluation.

Traps

  • Recognition and synthesis run in opposite directions.
  • A job accepted response is not completed analysis.
  • OCR text can contain hostile instructions.

Practise this topic

04 · RAG, Agents and Production Integration

Memory hook: Authorise evidence before retrieval, and authorise actions before tool execution.

Must remember

  • Build ingestion with source permissions, stable IDs, text/layout extraction, chunking, embeddings and index updates/deletes. Azure AI Search supports lexical, vector and hybrid retrieval with supported semantic ranking/enrichment capabilities. Vector similarity, ranking and grounded answer quality are different measurements.
  • Enforce tenant/document permissions using trusted identity context before evidence enters the model. Preserve source references for citations and freshness. Tune chunk size, overlap, filters and reranking on held-out questions; more retrieved text is not always better.
  • Agents need explicit roles/goals, conversation state, memory, typed tool schemas and limits. Integrate APIs, search, knowledge stores, custom functions and content-analysis tools through supported interfaces. A tool result is untrusted input and may itself contain injection.
  • Multi-agent systems require ownership of tasks/shared state, clear handoffs, failure propagation and end-to-end evaluation. Supervisor/delegation patterns add cost and latency. Prefer a deterministic workflow where known branches are sufficient; use rules for exact constraints rather than relying on model obedience.
  • Record project/deployment configuration, prompts, index versions and tool contracts in CI/CD. Test with managed identities/private networking and the actual application roles. Treat model, prompt and retrieval changes as releases with canary/rollback and compatibility checks.
  • Monitor tokens, quotas, time to first useful output, completion latency, tool failures and cost per successful task. Bound retry/reflection loops, validate structured output, cache with identity/freshness context and evaluate smaller-model routing where quality permits.

Choose under exam pressure

Requirement Choice and reason
Current internal knowledge with citations RAG with ACL-aware retrieval and support checks.
Several tools need controlled sequencing A workflow or bounded agent orchestration.
A model requests an irreversible action Validate policy and required approval before execution.

Traps

  • Reflection can repeat an error rather than correct it.
  • A cache without tenant context can leak data.
  • Changing an embedding model can require index migration.

Practise this topic

05 · AI Evaluation, Safety and Observability

Memory hook: Measure groundedness and task success alongside latency, cost and harm.

Must remember

  • Keep representative evaluation datasets separate from tuning examples. Include adversarial, no-answer, multi-turn and subgroup cases. Use rubrics for relevance, correctness, groundedness, safety and action success; calibrate model judges with human review.
  • Inspect retrieval/index health and ingestion freshness separately from model performance. Drift can affect input distributions, tool contracts or knowledge availability. A fluent answer with a real citation may still be unsupported by that source.
  • Configure supported safety filters/guardrails, risk detection and moderation. Test both direct injection and indirect instructions in retrieved documents, images and tool outputs. Restrict agent tools/credentials and enforce oversight outside the model.
  • Use managed identity, least-privilege roles, private networking where appropriate and controlled secret access. Redact sensitive telemetry; record provenance, approvals, model/prompt/tool versions and correlation IDs for investigation.
  • Trace each retrieval, model and tool step. Distinguish quota/rate limiting, invalid inputs, failed authorisation, network timeouts and model-quality failures. Retry only appropriate failures with bounded backoff; reconcile ambiguous external actions by operation ID.
  • Evaluate image/audio safety and accessibility as well as text. Human escalation, documented limitations and a feedback process remain necessary for high-consequence use. No model-generated confidence value is an independent guarantee.

Choose under exam pressure

Requirement Choice and reason
Quality drops after a deployment Compare fixed evaluation cases and component versions.
Agent repeats a tool indefinitely Enforce limits and inspect state/termination conditions.
Telemetry contains personal prompts Apply minimisation, redaction, access and retention policy.

Traps

  • A judge model has its own biases.
  • A passing safety filter is not proof of factual correctness.
  • More logging can increase privacy exposure.

Practise this topic

06 · Multimodal Generation and Extraction Pipelines

Memory hook: Keep the original evidence, its location and the requested transformation distinct.

Must remember

  • Image/video generation can use prompts and reference media with model-specific controls. Inpainting uses a mask to identify editable regions; other editing workflows change selected content or temporal segments. Validate model capabilities and output constraints rather than assuming every endpoint supports every edit.
  • Visual understanding includes captions, image questions, objects/regions and video-segment interpretation. Accessibility alt text should communicate relevant visible content concisely; extended descriptions serve a different purpose. Avoid inventing unseen details or treating uncertain classifications as facts.
  • Content Understanding analyzers can produce structured or Markdown representations for downstream use. Single-task and pro-mode processing have different capabilities/operating characteristics; choose using task complexity, supported modalities, latency and cost. Preserve layout, bounding/location information and timestamps where needed.
  • Text pipelines can extract entities/topics, summarise, classify tone/sentiment, detect sensitive content and translate with supported tools/models. Speech pipelines handle recognition, synthesis, translation and supported custom models. Test accents, noise, languages and domain vocabulary.
  • Retrieval ingestion may combine OCR, layout analysis, built-in/custom enrichment skills and vector/hybrid indexes. Tables, figures and document hierarchy need special handling; flattening everything into arbitrary text chunks can lose relationships.
  • Apply visual/content policy, supported watermarks and brand restrictions where required. Embedded text can inject instructions; generated media can contain unsafe content. Validate extracted fields against schema/business rules and route uncertain or consequential results for review.

Choose under exam pressure

Requirement Choice and reason
Change only part of an image A supported mask-based editing/inpainting workflow.
Extract fields and layout for RAG Content Understanding/OCR-layout pipeline with provenance.
Improve domain speech recognition Evaluate supported custom speech and representative recordings.

Traps

  • Image input support does not imply video generation support.
  • A caption is not a complete structured extraction.
  • A low-confidence field must not silently become authoritative business data.

Practise this topic

Search across every published topic.