| ISSUE 02/30 · COMPUTER SCIENCE |
~3 MIN |
NOW / AI INFRASTRUCTURE · WEEK 01 — HIDDEN RULES
Why AI systems cache your prompt
START HERE
When your phone shows the same photo twice, it can keep a nearby copy instead of downloading the file again. The instructions and question you give an AI are called a prompt. A prompt cache is saved computer work from reading a part of that prompt that has not changed. Without that saved work, the AI may spend extra time and electricity rereading the same beginning on each new model request.
|
/ WHY NOW
A coding assistant may send many model requests containing the same project instructions, tool descriptions, and earlier conversation. If every new request recomputes that unchanged beginning, the system spends extra time, electricity, and money. That is why prompt caching is in current AI infrastructure news: it turns an exact repeated beginning into reusable work.
|
/ THE IDEA
A cache is a nearby copy of something expensive to recreate. Your browser caches images. An AI prompt cache keeps the numerical work already done while reading an unchanged beginning, or prefix, of the prompt. Prefix matters. If the stable instructions stay at the beginning and new material is appended after them, the system can reuse the saved beginning. Change an early line and much of the later calculation may need to be rebuilt.
THE FORMAL IDEA
saved work ≈ repeated prefix × cache-hit rate
| repeated prefix = the unchanged text at the start | | cache-hit rate = how often a usable saved copy is found | | ≈ means this is a planning model, not a billing equation |
|
RUN THE TINY EXAMPLE
A project brief split into text pieces
A token is one small piece of text; stable brief = 20,000 tokens Request 1: read the brief + a 1,000-token question Request 2 with a cache hit: reuse the brief’s saved work + read a different 1,000-token question Move one instruction near the top: the reusable prefix may shrink
|
The answer is not copied from the cache. Only the work of reading the unchanged beginning is reused; the model still computes the new continuation.
/ SO WHAT?
This gives you a useful design rule for AI products: put stable instructions first, keep their wording and order steady, then append changing user data. The same idea explains why tidy agent context can be faster than a giant, constantly rewritten prompt.
ONE CAVEAT |
| A prompt cache is not long-term memory. Entries expire, providers implement them differently, and a hit saves repeated computation without guaranteeing the same answer. |
KEEP THIS
When the beginning repeats exactly, an AI system can reuse the reading and spend its effort on what changed.
|
NEXT: Post-quantum security starts before quantum computers
|