Prompt caching
Prompt caching lets the API reuse long context between calls, cutting cost up to 90% and latency up to 85% on long prompts.
Developers mark large static context (a book, codebase, instructions) to be cached. Cache reads cost about a tenth of base input price. For Claude 3.5 Sonnet that is $3.75 per M tokens to write and $0.30 to read. In one example, chatting with a 100K-token book saw 79% lower latency and 90% lower cost.
- Date
- Wednesday, 14 August 2024
- Lab
- Anthropic
- Kind
- feature
- Access
- closed API
- Price
- Claude 3.5 Sonnet cache write $3.75 / read $0.30 per M tokens
Figures
| Measure | Value | Measured by |
|---|---|---|
| Cost reduction (max) | up to 90% long prompts | company |
| Latency reduction (max) | up to 85% long prompts; 79% for a 100K-token book chat | company |
The captured blog metadata is garbled (shows a 2025 publication year alongside an 'updated 2024-12-17' line). 2024-08-14 is the public-beta launch date as best known; it was not confirmed from the captured page, which was later updated to describe general availability.
Sources
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check