Contextual Retrieval
Contextual Retrieval prepends chunk-specific context before embedding, cutting failed retrievals 35% (embeddings), 49% (with BM25) and 67% (with reranking).
A RAG technique in which Claude writes a short explanatory context for each document chunk before it is embedded and indexed with BM25. Generating the contextualized chunks costs about $1.02 per million document tokens when prompt caching is used.
- Date
- Thursday, 19 September 2024
- Lab
- Anthropic
- Kind
- feature
- Access
- paper only
- Price
- About $1.02 per M document tokens to contextualize chunks (with prompt caching)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Retrieval failure reduction, contextual embeddings | 35% vs standard embeddings | company |
| Retrieval failure reduction, + contextual BM25 | 49% | company |
| Retrieval failure reduction, + reranking | 67% | company |
Engineering blog technique, not a model release. Numbers are Anthropic's own tests.
Sources
This record was checked against its sources on 6 October 2026. How we check