Memory hook: Keep vectors aligned with the source.
Must remember
- Select external models by modality, language, quality, output format, dimensions, security and cost. Configure supported model endpoints and credentials/managed identities without embedding secrets in SQL text.
- Choose source columns and chunk boundaries according to meaning and update patterns. Maintain embeddings through supported triggers, Change Tracking/CDC, Functions, Logic Apps or Foundry workflows; track model/version alongside each vector.
- Full-text search matches terms; vector search finds semantic neighbors; hybrid search combines both. KNN can provide exact nearest-neighbor results; ANN trades some recall for speed/scale.
- Use compatible vector types, dimensions, indexes and distance metrics. Supported VECTOR_DISTANCE, VECTOR_SEARCH, normalization and property functions have distinct purposes; check the platform/version syntax.
- Reciprocal rank fusion combines ranked lists without assuming raw lexical and vector scores share a scale. Evaluate recall, relevance, latency and filtering together.
- RAG retrieves authorized context, serializes suitable structured data as JSON, calls the model through supported mechanisms such as sp_invoke_external_rest_endpoint, then validates/cites the answer. Retrieved text must not gain permission to execute arbitrary SQL.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Need exact terminology plus semantic similarity | Hybrid search with a measured ranking/fusion strategy. |
| Source documents changed after embedding | Update the affected chunks/vectors and remove stale entries. |
Traps
- Vectors from incompatible embedding models should not be compared as if they share one space.
- An ANN index does not guarantee the exact nearest result.
Active recall
1. Why store embedding model/version?
To detect incompatible vectors and support controlled re-embedding.
2. What does RRF combine?
Rankings from separate retrieval methods using rank positions.
3. When is exact KNN useful?
When exact recall is required and the candidate set/cost is manageable.
4. Why enforce authorization before generation?
Once unauthorized data enters the prompt, output filtering is an inadequate boundary.
5. What makes a RAG response trustworthy?
Relevant authorized evidence, faithful synthesis, citations and validation—not fluency alone.