Reviewed 10 October 2026 · Generative AI Developer Professional
Memory hook: Authorize evidence, bound actions, measure the whole task.
Model/data integration 31%; implementation 26%; safety/security 20%; efficiency 12%; testing 11%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — Analyze requirements and design GenAI solutions
- Translate outcomes into quality, modality, latency, availability, privacy, cost and integration constraints. Choose a deterministic workflow when it suffices; use RAG for external evidence and agents for justified dynamic actions. Define acceptance data, escalation behavior and regional/model availability before implementation.
1.2 — Select and configure FMs
- Compare models on representative tasks, supported input/output, context/output limits, licenses, latency, throughput and total cost. Tune sampling deliberately; evaluate smaller/routed models. Managed Bedrock access and custom SageMaker hosting imply different control and operational responsibilities.
1.3 — Implement data validation and processing pipelines for FM consumption
- Profile source ownership, rights, freshness, formats and tenant boundaries. Clean/deduplicate, extract structure, chunk semantically and preserve stable IDs/citations. Apply quality checks and deletion/update propagation; source deletion without index/cache cleanup leaves stale or forbidden evidence.
1.4 — Design and implement vector store solutions
- Choose vector stores by scale, latency, durability, filtering, tenancy and cost. Embedding model/dimensions must match index/query behavior; migrate and reevaluate when changing them. Approximate nearest-neighbor search trades retrieval recall for speed; encryption and authorization remain independent requirements.
1.5 — Design retrieval mechanisms for FM augmentation
- Compare keyword, vector and hybrid retrieval; metadata/ACL filters enforce scope, reranking improves order and adds cost. Measure retrieved-document relevance/recall separately from answer groundedness and citation support. Handle missing, stale and contradictory evidence with a supported no-answer path.
1.6 — Implement prompt engineering strategies and governance for FM interactions
- Version templates, instructions, examples, schemas and retrieval settings. Budget context for rules, input, conversation, evidence, tool output and response. Test few-shot/decomposition strategies on held-out tasks. Treat untrusted content as data; prompts cannot replace enforceable access policy.
2.1 — Implement agentic AI solutions and tool integrations
- Agents need typed tools, scoped identities, validation, durable state and bounded iterations. Bedrock Agents/AgentCore or frameworks such as Strands still require application controls. Separate planning from action authorization; use idempotency and appropriate human approval. Multi-agent handoffs need clear state ownership.
2.2 — Implement model deployment strategies
- Select managed model invocation, eligible provisioned throughput or custom hosting from traffic and control needs. Package dependencies, versions and configuration reproducibly. Use staged rollout with quality, safety, latency and cost gates; retain a compatible fallback model or release.
2.3 — Design and implement enterprise integration architectures
- Integrate enterprises through authenticated APIs, queues, Lambda and Step Functions as appropriate. Asynchronous jobs return a correlation ID and expose observable completion. Validate schemas, secrets, routes and identity propagation; do not give an agent unrestricted backend access.
2.4 — Implement FM API integrations
- Bedrock Converse standardizes supported conversational calls; model-specific APIs expose different capabilities. Validate fields/model support and handle streaming failures. Set timeouts, cancellation, bounded jittered retries and circuit breakers; distinguish quota throttling from invalid input or denied access.
2.5 — Implement application integration patterns and development tools
- Use modular retrieval, generation, tools and evaluation components with explicit contracts. AI-assisted coding tools can accelerate implementation but generated code needs review and tests. Cache/reuse only with correct authorization/freshness boundaries; trace each component rather than only the final response.
3.1 — Implement input and output safety controls
- Apply input filters, supported Guardrails, structured-output/schema checks and post-processing defense in depth. Test direct/indirect injection, jailbreaks, harmful content and unsafe actions. A valid JSON response can still contain false facts; grounding and business-rule validation are separate checks.
3.2 — Implement data security and privacy controls
- Enforce document/row/tenant access before adding evidence to context. Use IAM, KMS, TLS, private endpoints where needed and deliberate log redaction. Detect/mask PII with appropriate services, minimize collection and propagate retention/deletion to vectors, caches, prompts and agent memory.
3.3 — Implement AI governance and compliance mechanisms
- Track model/prompt/index versions, provenance, licenses, approvals and evaluation evidence. CloudTrail audits supported API actions; model invocation logging serves a different purpose and needs privacy controls. Record owner, risk acceptance, allowed uses and review triggers for provider/model changes.
3.4 — Implement responsible AI principles
- Measure fairness, safety and subgroup effects; explain limitations and evidence without treating generated reasoning as proof. Calibrate human/LLM judges and document model cards. Provide uncertainty/no-answer handling, contestability and meaningful oversight for consequential use.
4.1 — Implement cost optimization and resource efficiency strategies
- Optimize cost per successful task, including tokens, embeddings, vector queries, tools, storage and retries. Use model routing, shorter context, supported prompt caching and batch processing where quality/latency permit. Cache keys must respect caller access and source freshness; commitments need measured utilization.
4.2 — Optimize application performance
- Measure first-useful-output and total task latency separately. Profile retrieval, reranking, prompt assembly, generation and tools. Use supported streaming, safe parallelism, index/query optimization and appropriate token/output limits. Faster retrieval cannot fix an expensive serial chain of tool calls.
4.3 — Implement monitoring systems for GenAI applications
- Trace retrieval/model/tool spans with correlation IDs, versions, redacted errors and token counts. Alert on task failures, tail latency, throttles, loop rate, safety interventions and spending. Monitor vector freshness and tool outcomes; fluent responses can conceal failed external actions.
5.1 — Implement evaluation systems for GenAI
- Maintain golden tasks, no-answer cases, adversarial cases, tenant isolation and multi-turn tests. Separate retrieval, answer and action metrics; combine contract checks, human evaluation and calibrated LLM judges. Canary model/prompt/index changes and gate rare high-consequence failures as well as averages.
5.2 — Troubleshoot GenAI applications
- Localize faults: ingestion/index/filter/ranking, context truncation, model request parameters, authorization, tool selection/arguments/execution or final interpretation. Compare reproducible versions and traces. Repair the failing stage before adding tokens, retries or a larger model.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Correct source never retrieved | Inspect ingestion, index compatibility, ACL filters and ranking. |
| Source retrieved, answer unsupported | Inspect context use, generation and grounding/citation evaluation. |
| Agent claims success but backend unchanged | Verify tool result/external state; do not trust prose. |
| Store prompts for audit | Apply classification, redaction, retention, access and deletion requirements. |
| Timeout may duplicate an order | Reconcile operation ID before retrying; use idempotent execution. |
Traps
- A semantic cache can leak data across tenants.
- Temperature zero is not a truth guarantee.
- LLM-as-a-judge inherits bias and must be calibrated.
- A model update can change behavior without application code changes.
Verification cues
- Follow one interaction ID through retrieval, model invocation and tool execution.
- Explain the exact caller identity used for evidence retrieval and each side-effecting tool.
- Compare the same golden tasks across old/new model, prompt and index versions; capture quality, safety, latency and cost.
Last-pass active recall
1. Why isolate retrieval evaluation?
Missing authorized evidence differs from incorrect generation over good evidence; the fixes differ.
2. What must change when embedding dimensions change?
Use a compatible index and re-embed/migrate data as required, then reevaluate retrieval.
3. Why can schema-valid JSON be unsafe?
Its values may be unauthorized, fabricated or invalid for business rules.
4. What bounds an agent loop?
Iteration/time/cost limits, exit conditions, tool policy and observable escalation.
5. Which measure combines economics and usefulness?
Cost per successfully completed task under required quality, safety and latency constraints.
Sources and version check
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Foundation Models and Generative AI
Memory hook: A model predicts plausible output; your application must establish whether it is useful and supported.
Must remember
- A foundation model is pretrained on broad data and can support multiple downstream tasks. Transformers underpin many language models; diffusion models are common for image generation. Multimodal models process or generate more than one modality.
- A token is a model-specific text or data unit, not necessarily a word. A context window limits what a request can include; input and output consume capacity and may have different prices. Longer prompts and answers can increase cost and latency.
- Embeddings represent content as vectors for similarity operations. Chunking divides source content into retrievable pieces. Embeddings do not encrypt text and do not themselves generate a final answer.
- Bedrock offers managed access to supported models and application capabilities. SageMaker AI/JumpStart support a more custom model-development and hosting path. Model choice depends on modality, languages, quality, context/output length, region, licensing, latency, privacy and cost.
- Context engineering selects and organises instructions, retrieved evidence, tool results, conversation state and memory within the available context. Dumping an entire document collection into every request is costly and can obscure relevant facts.
- Generation can hallucinate, vary between calls and be hard to explain. Lower temperature generally reduces sampling variability; it does not guarantee correct or identical answers. Evaluate with real tasks rather than choosing the largest model by default.
- Compare token-based on-demand use, provisioned throughput where supported, caching, batch options and custom-model infrastructure. Business value includes task completion, time saved, user satisfaction and cost per successful interaction, not just tokens per second.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Quickly call several supported foundation models | Bedrock managed model APIs. |
| Need custom training and hosting control | SageMaker AI, with additional operating responsibilities. |
| Relevant facts are missing from the prompt | Improve context/retrieval before buying a larger model. |
Traps
- An embedding is not a lossless copy or a privacy boundary.
- A large context window does not guarantee reliable use of every supplied fact.
- Fluent wording is not evidence of factual correctness.
02 · Prompt Engineering, RAG and Fine-Tuning
Memory hook: Prompt for instructions, retrieve for current evidence, tune for learned behaviour.
Must remember
- A good prompt defines the task, relevant context, output format and constraints. Zero-shot uses no demonstration; one-shot/few-shot include examples. Templates standardise inputs. Version prompts and evaluation sets so improvements can be compared and rolled back; Bedrock Prompt Management supports managed prompt versions.
- Structured reasoning prompts can help decompose a task, but an explanation is not proof of a model's internal process or correctness. Prefer checkable intermediate results and concise justifications. Negative instructions alone are weak enforcement.
- RAG retrieves relevant content, adds it to a prompt and generates a grounded response. A typical path is ingest → clean → chunk → embed → index → retrieve → generate → cite. Bedrock Knowledge Bases manages supported parts of this workflow.
- Vector storage can use supported OpenSearch, Aurora/PostgreSQL or other integrations; the exact service feature matters. Hybrid keyword/vector retrieval, metadata filtering and reranking can improve results. Enforce user access before evidence enters the prompt.
- Fine-tuning changes model weights using task/domain examples; it can improve style and behaviour but is not a live fact lookup. Continued pretraining adapts using additional domain text. Instruction tuning uses instruction/response examples. RLHF uses human preference feedback. Distillation trains a smaller model from a larger model's outputs or behaviour.
- Curate representative, licensed, deduplicated data and hold out evaluation examples. Full pretraining has far higher data/compute requirements than prompt changes or retrieval. Prompt caching can reuse eligible repeated context; it does not repair stale source facts.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Answer questions about policies updated daily | RAG over authorised current documents. |
| Consistent specialised output style across many tasks | Evaluate prompting, then fine-tuning if needed. |
| Reduce a successful model's serving cost | Evaluate a smaller model or distillation against quality targets. |
Traps
- RAG does not alter model weights.
- Fine-tuning does not automatically grant access to new documents.
- A retrieved document can contain hostile instructions; treat it as untrusted data.
03 · Agents, Tools and AWS AI Platforms
Memory hook: The model proposes; tools act; policy decides whether an action is allowed.
Must remember
- An agent uses a model, instructions, state and tools to work toward a goal. A fixed workflow follows predefined steps; an agent can choose among actions dynamically. Use the simpler workflow when its decision structure is sufficient.
- A tool has an interface/schema, permissions and observable results. Validate arguments and outputs. Require human approval for consequential actions where appropriate. Retries need idempotency so a timeout does not create duplicate purchases or updates.
- MCP standardises how supported clients connect to tools and resources; it does not grant trust or make every exposed tool safe. Multi-agent designs may use a supervisor, delegation or peer interaction. More agents introduce coordination, latency and failure modes, not guaranteed quality.
- Short-term state tracks the current task; longer-term memory persists selected facts. Apply data minimisation, isolation, retention and deletion policies. Never mix one customer's retrieved data or memory with another's context.
- Bedrock Agents provides managed agent capabilities; AgentCore provides services for operating agents, including runtime and identity-related capabilities. Strands Agents is an agent-development framework. Choose components based on control and operating requirements, not similar names.
- Current AWS objectives also mention Amazon Quick for business AI experiences and Kiro for AI-assisted development. These are not substitutes for the underlying identity, model evaluation and application security controls. Product branding evolves; use the exam's current terminology.
- Evaluate whole-task completion, valid tool choice, argument accuracy, loop rate, latency and total cost. Add execution limits, safe failure paths and traces so an agent that repeats a tool indefinitely can be diagnosed and stopped.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Known approval steps and deterministic branches | Workflow orchestration. |
| A model must choose among approved business tools | An agent with scoped permissions and validation. |
| Several specialised agents share work | Explicit orchestration, state boundaries and end-to-end evaluation. |
Traps
- Tool discovery is not tool authorisation.
- Memory can preserve incorrect or sensitive information.
- Successful text generation is not the same as successful completion of an external action.
04 · Evaluation and Responsible AI
Memory hook: Measure the answer, the experience and the harm; averages can hide who fails.
Must remember
- Use representative held-out tasks, stable baselines and subgroup analysis. Compare quality, robustness, safety, latency, cost and business outcomes. A benchmark unrelated to the real workflow is weak evidence of suitability.
- BLEU emphasises n-gram precision against references and is associated with translation. ROUGE uses overlap/recall-oriented measures often applied to summaries. BERTScore uses contextual representations for semantic similarity. None alone proves factual correctness or useful business outcomes.
- LLM-as-a-judge can scale evaluation but introduces judge bias, model/version sensitivity and prompt dependence. Calibrate against human judgement and inspect disagreements. Bedrock evaluation capabilities support supported automated and human evaluation approaches.
- Evaluate RAG retrieval separately from answer groundedness and relevance. Evaluate an agent's tool use, action validity and completed task, not just its final wording. Include adversarial inputs and safe refusal/escalation tests.
- Responsible AI includes fairness, robustness, safety, privacy, transparency, explainability, accountability and veracity. Representative data, label review, human audits and subgroup metrics help find harmful differences hidden by aggregate accuracy.
- A transparent system exposes relevant workings and limitations; an explanation helps a person understand a particular result or behaviour. Model cards document intended use, evidence, limitations and risk. An open-source model is not automatically interpretable, safe or appropriately licensed.
- Consider intellectual-property rights, deceptive or biased outputs, environmental impact and user trust. Human-centred design needs clear AI disclosure, feedback and contestability where appropriate. Guardrails reduce specified risks but cannot certify that every response is harmless or true.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Summary wording differs but meaning is similar | Use semantic and human evaluation alongside overlap metrics. |
| High overall accuracy but poor outcomes for one group | Subgroup analysis and fairness investigation. |
| High-consequence decision | Human oversight, documented limits and a suitable error/appeal process. |
Traps
- Fairness has multiple definitions and trade-offs.
- A hallucination score is evidence, not a guarantee.
- Removing all sensitive columns does not necessarily remove proxy bias.
05 · AI Security, Privacy and Governance
Memory hook: Protect the data path and the action path, then keep evidence of both.
Must remember
- Apply least-privilege IAM roles to models, tools, data stores and logs. Separate end-user identity from workload identity. Agent identity and policy features support controls, but the application still needs tenant isolation and scoped tool permissions.
- Use TLS in transit, suitable encryption at rest and controlled KMS key access. PrivateLink provides supported private connectivity; it does not replace identity authorisation. Macie helps discover sensitive S3 data. Secrets should not be included in prompts or source code.
- Prompt injection tries to turn untrusted text into instructions. It may arrive in a user message, retrieved page or tool result. Poisoning corrupts training or indexed data. Jailbreaking seeks to defeat safety behaviour. Validate inputs/outputs, restrict actions and treat retrieved content as data.
- Bedrock Guardrails can apply configured content, topic, sensitive-information and other supported controls. Grounding and output validation can help detect unsupported statements. Do not use model self-reported confidence as the sole authority for high-risk decisions.
- Track provenance, licences and lineage from source data through transformations, model versions and outputs. Define residency, retention and deletion requirements for prompts, logs, embeddings and memory as well as primary datasets.
- CloudTrail supplies supported API audit events; Config evaluates resource configuration; Inspector assesses supported workload vulnerabilities; Artifact supplies AWS compliance evidence; Trusted Advisor highlights supported recommendations. Choose evidence according to the question.
- Governance needs accountable owners, approval gates, review cadence, staff training and documented exceptions. Use a risk framework appropriate to the application and shared-responsibility model. Service compliance does not certify that your own data collection or use is lawful.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A retrieved document tells an agent to reveal secrets | Treat as injection; enforce permissions and tool constraints outside the model. |
| Need evidence of configuration compliance | Config plus the relevant audit process. |
| Need to prove where training material came from | Data lineage, provenance and licensing records. |
Traps
- A private network does not prevent an authorised application from leaking data.
- Encrypting vectors does not resolve the right to retain their source data.
- Logging every prompt without redaction can create a second sensitive-data store.
06 · Production Retrieval and Context Engineering
Memory hook: Retrieve evidence the caller may use, measure whether it is relevant, and fit it into the context deliberately.
Must remember
- Profile source formats, ownership, licences, update/delete frequency, tenant boundaries and quality before ingestion. Extract text/structure, normalise encoding, remove duplicates and keep stable document/chunk IDs with provenance. Preserve page/section references for citations.
- Chunk by semantic boundaries where practical; tune size and overlap against retrieval quality, token cost and document structure. Tables and images may need multimodal or structure-aware processing. Re-embedding requires compatible model dimensions and a planned index migration.
- Select vector storage by scale, latency, filtering, durability, tenancy, updates and operations. Approximate nearest-neighbour search trades speed against recall; metadata filters and hybrid lexical/vector search solve different retrieval failures. Reranking can improve relevance at added latency/cost.
- Enforce source/document permissions during retrieval, not by asking the model to hide forbidden information afterwards. Tenant filters must come from trusted identity context rather than an unvalidated user argument. Propagate revocations and deletion into indexes and caches.
- Separate retrieval evaluation from generation: relevant-chunk recall, ranking quality, answer relevance, groundedness and citation correctness. A plausible citation can point to a document that does not support the claim. Handle no-answer cases and conflicting/stale evidence explicitly.
- Context budgets include instructions, conversation, retrieved chunks, tool results and output allowance. Summarise or select context with traceable rules; avoid truncating crucial constraints. Prompt templates, model/embedding versions and retrieval settings belong in release metadata.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Correct document exists but ranks too low | Inspect embedding/query representation, filters, hybrid search and reranking. |
| Only one tenant should see a document | Enforce trusted tenant/ACL filters before model context construction. |
| Index must change embedding dimensions | Build and evaluate a compatible new index, then migrate traffic. |
Traps
- Vector similarity is not factual entailment.
- Document deletion without index/cache invalidation can preserve access.
- A citation needs support validation, not just a valid URL.
07 · Agent Integration, Inference APIs and Operations
Memory hook: Treat every model call and tool call as a bounded distributed-system operation.
Must remember
- Use supported APIs such as Bedrock Converse when their common interface meets model needs; model-specific APIs may expose different features. Validate request/response schemas, supported parameters, model access and Region. Streaming changes response delivery and error handling, not the need for authorisation.
- Build timeouts, cancellation, backoff with jitter, retry budgets and circuit breakers. A rate-limit response differs from a malformed prompt or denied request. Check token/request quotas and actual model throughput. Retries after tool execution need idempotency and a recorded action outcome.
- Separate agent planning from tool execution policy. Use typed schemas, scoped credentials, allowlisted operations, parameter validation and approval gates for consequential actions. AgentCore/Bedrock agent capabilities or framework-based deployments still need observable orchestration and bounded iteration.
- Integrate business systems using appropriate APIs, Lambda, queues and Step Functions. Long-running work can return a job identifier and complete asynchronously. Events and callbacks need correlation, authentication and duplicate handling. Multi-agent systems need ownership of shared state and explicit handoff contracts.
- Optimise cost through model/task routing, smaller models, context reduction, supported prompt caching, batching and appropriate throughput commitment. Include retrieval, embeddings, storage, orchestration, tool calls and failed retries in cost per successful task. Cache by authorisation and freshness context, not only raw question text.
- Trace model, retrieval and tool spans with correlation IDs, versions, token counts and redacted errors. Monitor end-to-end completion, latency percentiles, throttling, safety interventions and cost. Protect captured prompts/responses under a deliberate privacy/retention policy.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Interactive output must start quickly | Use supported streaming and measure time to first useful output. |
| An expensive model handles simple classification | Evaluate routing to a smaller suitable model against quality targets. |
| A tool timed out after possibly committing a change | Reconcile by operation ID before retrying a side effect. |
Traps
- Faster first-token latency is not necessarily faster task completion.
- A cache can leak data if identity context is missing from its key/control.
- More retries can deepen a throttling cascade.
08 · GenAI Testing and Failure Diagnosis
Memory hook: Freeze the test conditions, separate failure stages and verify the final business action.
Must remember
- Maintain representative golden tasks, adversarial cases, no-answer cases, multi-turn interactions and tenant-boundary tests. Keep development and final evaluation sets separate. Record model, prompt, retrieval/index and tool versions to reproduce regressions.
- Use deterministic schema/contract checks where possible, semantic metrics where useful and calibrated human or model judgement for subjective outcomes. An LLM judge needs a rubric, bias checks and validation against human decisions.
- Decompose RAG failures into ingestion, retrieval, reranking, context assembly and answer generation. Decompose agent failures into planning, tool selection, argument formation, authorisation, execution and result interpretation. Fixing the wrong stage can increase cost without improving correctness.
- Test prompt injection through user input, retrieved documents and tool output. Check output encoding, command/query injection, sensitive-data leakage, refusal boundaries and side effects. Guardrails require configuration and regression tests; they are not a universal policy engine.
- Load-test realistic context sizes, concurrency, streaming and downstream dependencies. Monitor quota use and degradation paths. Canary a new prompt/model/index and retain an explicit rollback; a provider model update can change behaviour even if application code is unchanged.
- Evaluate against business acceptance criteria: task completion, correct external state, user satisfaction, safety, latency and total cost. Document residual limitations and keep a human escalation option for cases that cannot be reliably automated.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A new prompt improves averages but fails rare critical tasks | Gate on risk-weighted cases and subgroup/task categories. |
| Generated JSON is malformed | Enforce/validate supported structured output and handle validation failures. |
| Agent says it booked a meeting but calendar is empty | Verify tool outcome and external state, not prose. |
Traps
- A unit test with mocked model output cannot measure model quality.
- A benchmark score can hide unsafe tail cases.
- Evaluation datasets can become contaminated by repeated tuning.