Prompt Caching Explained: How to Cut AI API Costs Without Changing Your Model
Cached input tokens bill at a tenth of the standard rate. Here is exactly what gets cached, the 1,024-token minimum, why writes cost 1.25x, and the prompt-ordering mistakes that silently destroy your hit rate.