Memory hook: Keep the original evidence, its location and the requested transformation distinct.
Must remember
- Image/video generation can use prompts and reference media with model-specific controls. Inpainting uses a mask to identify editable regions; other editing workflows change selected content or temporal segments. Validate model capabilities and output constraints rather than assuming every endpoint supports every edit.
- Visual understanding includes captions, image questions, objects/regions and video-segment interpretation. Accessibility alt text should communicate relevant visible content concisely; extended descriptions serve a different purpose. Avoid inventing unseen details or treating uncertain classifications as facts.
- Content Understanding analyzers can produce structured or Markdown representations for downstream use. Single-task and pro-mode processing have different capabilities/operating characteristics; choose using task complexity, supported modalities, latency and cost. Preserve layout, bounding/location information and timestamps where needed.
- Text pipelines can extract entities/topics, summarise, classify tone/sentiment, detect sensitive content and translate with supported tools/models. Speech pipelines handle recognition, synthesis, translation and supported custom models. Test accents, noise, languages and domain vocabulary.
- Retrieval ingestion may combine OCR, layout analysis, built-in/custom enrichment skills and vector/hybrid indexes. Tables, figures and document hierarchy need special handling; flattening everything into arbitrary text chunks can lose relationships.
- Apply visual/content policy, supported watermarks and brand restrictions where required. Embedded text can inject instructions; generated media can contain unsafe content. Validate extracted fields against schema/business rules and route uncertain or consequential results for review.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Change only part of an image | A supported mask-based editing/inpainting workflow. |
| Extract fields and layout for RAG | Content Understanding/OCR-layout pipeline with provenance. |
| Improve domain speech recognition | Evaluate supported custom speech and representative recordings. |
Traps
- Image input support does not imply video generation support.
- A caption is not a complete structured extraction.
- A low-confidence field must not silently become authoritative business data.
Active recall
1. What does an inpainting mask specify?
The image region intended for editing under the selected workflow.
2. Why preserve layout in document extraction?
Relationships between headings, tables, labels and values can carry meaning.
3. Why test audio under realistic conditions?
Noise, accents and vocabulary materially affect recognition quality.
4. What is indirect visual prompt injection?
Instructions embedded in an image/document that attempt to override the application's trusted instructions.
5. What should downstream systems receive with extracted values?
A validated schema and useful source/provenance evidence, with uncertainty handled explicitly.