In the previous Session State Sync chapter, we learned how to keep our conversation history safe and consistent.
As our conversation history grows, we face a new problem: Size. Sending the entire history of a project (thousands of lines of code) to the AI for every single message is expensive and slow.
Imagine you visit a coffee shop every morning.
In the world of LLMs, Prompt Caching is the barista remembering your context. If the cache works, the API is fast and cheap (you pay 90% less!).
However, prompt caching is fragile. If you change a single character in your System Prompt or add a new Tool, the cache "breaks." The API treats you like a stranger, processes everything from scratch, and charges you full price.
The Prompt Cache Monitor (implemented in promptCacheBreakDetection.ts) is an intelligent observer. It watches every message you send.
If the cache breaks (latency spikes), it tells you exactly why:
You are a developer tuning the "System Prompt" (the instructions that tell Claude how to behave). You change "Be helpful" to "Be very helpful."
Suddenly, your API costs double. You don't know why. The Prompt Cache Monitor detects this and logs: [PROMPT CACHE BREAK] system prompt changed (+5 chars). Now you know the cause.
This system runs automatically in the background, but understanding how it works helps you write better code. It operates in two phases: Record (Before sending) and Check (After receiving).
Before we send a request to Claude, we take a "Snapshot" of everything that might affect the cache.
import { recordPromptState } from './promptCacheBreakDetection.js';
// Before calling the API...
recordPromptState({
model: 'claude-3-5-sonnet',
system: systemPrompts, // Your instructions
toolSchemas: tools, // Your available tools
querySource: 'repl_main' // Where is this coming from?
});
What happens here? The monitor calculates a "fingerprint" (hash) of your prompt and tools. It saves this fingerprint in memory.
After Claude replies, the API tells us how many tokens were "read from cache." We pass this number to the monitor.
import { checkResponseForCacheBreak } from './promptCacheBreakDetection.js';
// After getting the response...
await checkResponseForCacheBreak(
'repl_main', // The source key
cacheReadTokens, // e.g., 0 (Bad) or 5000 (Good)
createdTokens,
messages
);
What happens here?
If cacheReadTokens is significantly lower than the previous request, the monitor compares the current fingerprint with the old one and logs the difference.
How does the computer know why the cache broke? It uses a technique called Hashing and Diffing.
The magic happens in promptCacheBreakDetection.ts. Let's break down the logic.
We don't want to store huge strings of text in memory. Instead, we convert objects into a simple number (a hash). If the text changes, the number changes.
// inside promptCacheBreakDetection.ts
function computeHash(data: unknown): number {
// Convert object to string
const str = JSON.stringify(data);
// Create a unique number from that string
// If even one comma changes, this number changes completely
return djb2Hash(str);
}
We use a Map to remember the state of the previous successful request. This acts as our baseline.
const previousStateBySource = new Map<string, PreviousState>();
// When recording state...
const prev = previousStateBySource.get(key);
if (!prev) {
// First time? Just save the current state and exit.
previousStateBySource.set(key, {
systemHash: computeHash(system),
toolsHash: computeHash(tools),
// ... other fields
});
return;
}
When a new request comes in, we compare the new hash to the old hash. If they are different, we note it as a "Pending Change."
// Compare current vs previous
const systemPromptChanged = systemHash !== prev.systemHash;
const toolSchemasChanged = toolsHash !== prev.toolsHash;
if (systemPromptChanged || toolSchemasChanged) {
// We found a difference! Store this fact.
prev.pendingChanges = {
systemPromptChanged,
toolSchemasChanged,
// ... details about what changed
};
}
When the response comes back, we look at the token usage. If the cache usage dropped significantly, we check our "Pending Changes" list to explain why.
export async function checkResponseForCacheBreak(
cacheReadTokens,
prevCacheReadTokens
) {
// Did we lose a significant amount of cache?
const tokenDrop = prevCacheReadTokens - cacheReadTokens;
// If cache is healthy (95% hit), do nothing.
if (cacheReadTokens >= prevCacheReadTokens * 0.95) return;
// If cache broke, look at the notes we made earlier
if (pendingChanges.systemPromptChanged) {
console.warn("Cache Break Reason: System prompt changed");
}
}
Sometimes you didn't change anything, but the cache still broke. Why? Because caches expire (usually after 5 minutes or 1 hour of inactivity).
The monitor calculates the time gap to detect this.
const timeSinceLastMsg = Date.now() - lastMsgTime;
if (parts.length === 0) {
// If no code changed, check the clock
if (timeSinceLastMsg > 5 * 60 * 1000) {
reason = 'possible 5min TTL expiry';
}
}
In this chapter, we explored the Prompt Cache Monitor.
Now we have a fully functional system: Clients, Resiliency, Files, Sync, and Optimization. But how do we see all these metrics in one place?
Next Chapter: Telemetry & Observability
Generated by Code IQ