Memory hook: A model predicts plausible output; your application must establish whether it is useful and supported.
Must remember
- A foundation model is pretrained on broad data and can support multiple downstream tasks. Transformers underpin many language models; diffusion models are common for image generation. Multimodal models process or generate more than one modality.
- A token is a model-specific text or data unit, not necessarily a word. A context window limits what a request can include; input and output consume capacity and may have different prices. Longer prompts and answers can increase cost and latency.
- Embeddings represent content as vectors for similarity operations. Chunking divides source content into retrievable pieces. Embeddings do not encrypt text and do not themselves generate a final answer.
- Bedrock offers managed access to supported models and application capabilities. SageMaker AI/JumpStart support a more custom model-development and hosting path. Model choice depends on modality, languages, quality, context/output length, region, licensing, latency, privacy and cost.
- Context engineering selects and organises instructions, retrieved evidence, tool results, conversation state and memory within the available context. Dumping an entire document collection into every request is costly and can obscure relevant facts.
- Generation can hallucinate, vary between calls and be hard to explain. Lower temperature generally reduces sampling variability; it does not guarantee correct or identical answers. Evaluate with real tasks rather than choosing the largest model by default.
- Compare token-based on-demand use, provisioned throughput where supported, caching, batch options and custom-model infrastructure. Business value includes task completion, time saved, user satisfaction and cost per successful interaction, not just tokens per second.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Quickly call several supported foundation models | Bedrock managed model APIs. |
| Need custom training and hosting control | SageMaker AI, with additional operating responsibilities. |
| Relevant facts are missing from the prompt | Improve context/retrieval before buying a larger model. |
Traps
- An embedding is not a lossless copy or a privacy boundary.
- A large context window does not guarantee reliable use of every supplied fact.
- Fluent wording is not evidence of factual correctness.
Active recall
1. Are 1,000 words always 1,000 tokens?
No. Tokenisation depends on the model and content.
2. What separates a multimodal model from a text-only model?
Its supported input/output modalities, such as text plus images or audio.
3. Does temperature zero eliminate hallucination?
No. It changes sampling behaviour, not the factual basis of the response.
4. Why measure cost per successful task?
A cheaper call may require more retries or produce unusable output.
5. What does vector similarity tell you?
Relatedness under the embedding representation; it does not prove a retrieved claim is true.