Memory hook: Treat every model call and tool call as a bounded distributed-system operation.
Must remember
- Use supported APIs such as Bedrock Converse when their common interface meets model needs; model-specific APIs may expose different features. Validate request/response schemas, supported parameters, model access and Region. Streaming changes response delivery and error handling, not the need for authorisation.
- Build timeouts, cancellation, backoff with jitter, retry budgets and circuit breakers. A rate-limit response differs from a malformed prompt or denied request. Check token/request quotas and actual model throughput. Retries after tool execution need idempotency and a recorded action outcome.
- Separate agent planning from tool execution policy. Use typed schemas, scoped credentials, allowlisted operations, parameter validation and approval gates for consequential actions. AgentCore/Bedrock agent capabilities or framework-based deployments still need observable orchestration and bounded iteration.
- Integrate business systems using appropriate APIs, Lambda, queues and Step Functions. Long-running work can return a job identifier and complete asynchronously. Events and callbacks need correlation, authentication and duplicate handling. Multi-agent systems need ownership of shared state and explicit handoff contracts.
- Optimise cost through model/task routing, smaller models, context reduction, supported prompt caching, batching and appropriate throughput commitment. Include retrieval, embeddings, storage, orchestration, tool calls and failed retries in cost per successful task. Cache by authorisation and freshness context, not only raw question text.
- Trace model, retrieval and tool spans with correlation IDs, versions, token counts and redacted errors. Monitor end-to-end completion, latency percentiles, throttling, safety interventions and cost. Protect captured prompts/responses under a deliberate privacy/retention policy.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Interactive output must start quickly | Use supported streaming and measure time to first useful output. |
| An expensive model handles simple classification | Evaluate routing to a smaller suitable model against quality targets. |
| A tool timed out after possibly committing a change | Reconcile by operation ID before retrying a side effect. |
Traps
- Faster first-token latency is not necessarily faster task completion.
- A cache can leak data if identity context is missing from its key/control.
- More retries can deepen a throttling cascade.
Active recall
1. What should happen to a permanent validation error?
Return a useful controlled failure; do not retry it indefinitely.
2. Why trace tool calls separately?
They can dominate latency, fail independently or create side effects even when model generation succeeds.
3. What prevents an agent from looping without limit?
Iteration/time/cost bounds, state checks and a defined escalation/failure path.
4. Why include failed requests in cost metrics?
They consume resources without completing the user's goal.
5. What differs between API authentication and tool authorisation?
The first identifies the caller; the second permits a specific action on a specific resource.