Prompt Caching Explained: Cut Repeated Input Costs by Around 90%

Jonhisking · October 2, 2026

If your app sends the same long instructions, documents or tool definitions with every request, you are paying full price for text the model has already seen. Prompt caching fixes that. It is the single biggest cost saving available on most AI APIs, often cutting the price of repeated input by around 90%, and it needs only a small change in how you build your prompts.

What prompt caching does

When a model processes a prompt, it does a lot of computation on every token before it starts writing. If the next request begins with exactly the same text, the provider can keep the result of that computation and reuse it instead of doing it again.

Providers pass that saving on: tokens read from the cache are billed at a fraction of the normal input price. Discounts of about 90% (paying around 10% of the normal price) are common, though the exact rate depends on the provider and model.

The one rule: the cached part must be an identical prefix

Caching works on the beginning of the prompt. The provider looks for the longest stretch from the very first token that exactly matches a recent request. As soon as one token differs, everything after it is processed and billed normally.

That means order matters:

Cache-friendly — fixed content first, changing content last:

  1. System instructions (never change)
  2. Tool definitions (never change)
  3. Reference documents (change rarely)
  4. Conversation history (grows, but the earlier part stays the same)
  5. The user's new message (always new)

Cache-breaking — a timestamp, user name or random ID at the top of the system prompt. Because the first line differs every time, nothing after it can be reused.

A worked example

A support assistant has 10,000 tokens of fixed instructions and product documentation, followed by a short question. It answers 1,000 questions an hour. Input price: $2 per million tokens (GPT-6 Sol, checked October 1, 2026).

Fixed part billed Cost per hour
Without caching 1,000 × 10,000 tokens at $2/M $20.00
With caching (90% off cached reads) 1 full read + 999 cached reads at $0.20/M about $2.02

Over a month of steady traffic, that is the difference between roughly $14,600 and $1,500 for the fixed part of the prompt alone.

Details that vary by provider

Caching works differently on each platform, so check the documentation for the one you use. Things to look for:

Where caching helps most

Where it does not help

A checklist

  1. Move everything that changes per request (dates, names, IDs, the user's message) to the end of the prompt.
  2. Keep system prompts and tool definitions byte-for-byte identical across requests. Even a changed space breaks the match.
  3. Put shared documents before per-user content.
  4. Check your API responses: most providers report how many input tokens were read from the cache. If that number is zero, something at the top of your prompt is changing.
  5. For low-traffic apps, see whether a longer cache lifetime is available and worth it.

Related

Measure your system prompt and see what it costs per request on each model.

Open the token counter →