AI Research Atlas

Prompt caching

Anthropic · 14 August 2024

Prompt caching lets the API reuse long context between calls, cutting cost up to 90% and latency up to 85% on long prompts.

Developers mark large static context (a book, codebase, instructions) to be cached. Cache reads cost about a tenth of base input price. For Claude 3.5 Sonnet that is $3.75 per M tokens to write and $0.30 to read. In one example, chatting with a 100K-token book saw 79% lower latency and 90% lower cost.

Date
Wednesday, 14 August 2024
Lab
Anthropic
Kind
feature
Access
closed API
Price
Claude 3.5 Sonnet cache write $3.75 / read $0.30 per M tokens

Figures

MeasureValueMeasured by
Cost reduction (max)up to 90%
long prompts
company
Latency reduction (max)up to 85%
long prompts; 79% for a 100K-token book chat
company

The captured blog metadata is garbled (shows a 2025 publication year alongside an 'updated 2024-12-17' line). 2024-08-14 is the public-beta launch date as best known; it was not confirmed from the captured page, which was later updated to describe general availability.

Sources

  1. claude.com/blog/prompt-caching

This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check

Related