Welcome to the final chapter of our btw tutorial series!
In the previous chapter, Async Execution & Abort Control, we learned how to manage the lifecycle of our AI request so the application remains responsive and crash-free.
Now, we tackle the final challenge: Memory and Performance.
When you ask a side question, the AI needs to know what you were working on (Context). However, sending that data can be slow and expensive. Furthermore, we don't want our "side question" to clutter up the main chat history.
In this chapter, we will implement Context Forking & Caching.
Imagine you are writing a very long essay on a physical piece of paper. You suddenly need to do some quick math to check a fact.
Option A (The Bad Way): You write the math equation directly on your essay. Now your essay is messy, and you have to erase it later.
Option B (The btw Way):
You take your essay to a photocopier. You make a copy. You scribble your math on the copy. Once you get the answer, you throw the copy away and go back to the pristine original.
This is Context Forking.
Now, imagine the photocopier charges you $1.00 for every page it scans. If your essay is 100 pages long, that's expensive!
But, if the photocopier sees that the first 99 pages are exactly the same as the last time you used it, it skips scanning them and only charges you for the new page.
This is Prompt Caching. To use it, we must ensure the data we send is byte-identical to the previous request.
This is a bundle of data that represents the "State of the World" before you asked your side question. It includes:
getLastCacheSafeParams)The main application stores a "snapshot" of the last thing it sent to the AI. If we can grab this snapshot, we guarantee that our data matches the cache exactly.
We will use a function called buildCacheSafeParams. This acts as our "Smart Photocopier."
Sometimes, you might type btw while the main AI is still typing a sentence. We don't want to send a half-finished sentence to our fork, as it might confuse the AI.
// Helper function to remove half-finished messages
function stripInProgressAssistantMessage(messages: Message[]) {
const last = messages.at(-1);
// If the last message was from the AI and wasn't finished...
if (last?.type === 'assistant' && last.message.stop_reason === null) {
// ...chop it off!
return messages.slice(0, -1);
}
return messages;
}
Now for the core logic. We check if the main application has a saved snapshot.
// btw.tsx
async function buildCacheSafeParams(context) {
// 1. Clean up the current messages
const cleanMessages = stripInProgressAssistantMessage(context.messages);
// 2. Ask: "Do we have a cached snapshot?"
const saved = getLastCacheSafeParams();
If saved exists, we reuse its heavy data (systemPrompt, userContext). This ensures the "prefix" of our request is identical to what the API has seen before, triggering the cache optimization.
if (saved) {
return {
// REUSE the exact objects from memory
systemPrompt: saved.systemPrompt,
userContext: saved.userContext,
systemContext: saved.systemContext,
// Add current tool context
toolUseContext: context,
forkContextMessages: cleanMessages
};
}
If this is the very first thing the user has done, there might not be a snapshot yet. In that case, we have to do the hard work of building the parameters from scratch.
// If no cache, we must build it manually (Slower)
const [rawSystemPrompt, userContext, systemContext] = await Promise.all([
getSystemPrompt(context.options.tools, ...),
getUserContext(),
getSystemContext()
]);
return {
systemPrompt: asSystemPrompt(rawSystemPrompt),
userContext,
systemContext,
toolUseContext: context,
forkContextMessages: cleanMessages
};
} // end function
How does this flow work when the user actually runs the command?
btw Command: User types /btw.btw grabs the parameters used for Request A.
Finally, let's look at how this fits into our useEffect in btw.tsx. This connects the logic we just wrote to the UI we built in Chapter 2 and the Async handler from Chapter 4.
// btw.tsx - inside BtwSideQuestion component
useEffect(() => {
async function fetchResponse() {
try {
// 1. Fork the context efficiently
const cacheSafeParams = await buildCacheSafeParams(context);
// 2. Run the question using the forked context
const result = await runSideQuestion({
question,
cacheSafeParams
});
// ... handle result
} catch (err) { /* handle error */ }
}
// ...
}, [question]);
Congratulations! You have successfully built the btw command from scratch.
Let's review what we built across these five chapters:
You now have a fully functional, high-performance CLI tool that integrates seamlessly with an AI agent. By using these patternsβLazy Loading, Component State, AbortControllers, and Cache Forkingβyou can build robust terminal applications that feel professional and snappy.
Happy Coding!
Generated by Code IQ