In the previous chapter, Incremental Context Cursor, we taught the system to efficiently track which messages need to be analyzed. We know what to read. Now we must decide how to read it without annoying the user.
Imagine you are chatting with an AI. You say something simple like, "Hello." If the AI decides to save a memory right then, and it runs in the main process, you have to wait.
This delay breaks the flow of conversation. We want the memory system to be invisible.
To solve this, we use Forked Agent Execution.
Think of a courtroom trial.
Periodically, the stenographer quietly slips out the back door to file a report (write a memory) based on what just happened. The Judge doesn't stop the trial. The Judge keeps talking to you while the Stenographer is working in another room.
Scenario:
Goal: The AI answers about the weather immediately. In the background, a second, invisible AI process analyzes the project requirement and saves it to a file.
In software terms, a "fork" splits a process into two.
Running two agents sounds expensive, right? This system uses a smart optimization called Prompt Caching. Because the Main Agent and the Forked Agent share the exact same chat history (up until the fork happens), the AI provider (like Anthropic) reuses the "processing work" it already did.
This makes the background agent extremely fast and much cheaper than a fresh start.
How does the code handle this split? It uses a "Fire-and-Forget" pattern.
When the main agent finishes a turn (stops generating text), the system calls executeExtractMemories.
// extractMemories.ts - Public Entry Point
export async function executeExtractMemories(
context: REPLHookContext,
): Promise<void> {
// We call the internal extractor, but we don't 'await' it blocking the user.
// We let it run in the background.
await extractor?.(context)
}
Explanation: This function kicks off the process. It doesn't return any text to the user. It just signals the background worker to start.
Inside the background process, we use runForkedAgent. This is a special utility that creates our "Stenographer."
// extractMemories.ts - Inside runExtraction()
const result = await runForkedAgent({
// The 'userPrompt' contains our instructions from Chapter 1
promptMessages: [createUserMessage({ content: userPrompt })],
// We share the cache params so this is cheap and fast
cacheSafeParams: cacheSafeParams,
// Crucial: This agent is invisible. It does not write to the main transcript.
skipTranscript: true,
// Limit the agent so it doesn't get stuck in a loop
maxTurns: 5,
})
Explanation:
promptMessages: This is where we pass the "Employee Handbook" (Chapter 1) and the File List (Chapter 2).skipTranscript: true: This ensures the background agent's internal thoughts ("I should save this file...") never appear in your chat window.What if you close the application while the background agent is still writing a file? We don't want to corrupt data. We use a "Drain" function to ensure pending work finishes before the app exits.
// extractMemories.ts
export async function drainPendingExtraction(timeoutMs?: number): Promise<void> {
// Wait for all background tasks (inFlightExtractions) to finish
await drainer(timeoutMs)
}
Here is how the Main conversation and the Memory extraction run side-by-side.
To manage this background state, we use a coding pattern called a Closure. We wrap all our state variables inside an init function. This ensures that if we run multiple tests or sessions, the variables don't leak into each other.
// extractMemories.ts
export function initExtractMemories(): void {
// 1. Define state variables inside the function scope
const inFlightExtractions = new Set<Promise<void>>()
let inProgress = false
// 2. Define the worker function
async function runExtraction(inputs) {
inProgress = true
try {
// ... perform the work ...
} finally {
inProgress = false
}
}
// 3. Assign to the global export so the outside world can call it
extractor = async (context) => {
const p = runExtraction({ context })
inFlightExtractions.add(p) // Track the promise
await p
inFlightExtractions.delete(p)
}
}
Explanation:
inProgress: Acts as a lock. If the agent is already working, we don't start a second parallel worker (that would be confusing!).inFlightExtractions: A list of active jobs. The "Drain" function checks this list to know when it's safe to quit.The Forked Agent Execution model is the secret sauce that makes memory extraction feel magical.
skipTranscript).Now we have a background agent that is ready to write files. But waitβthis background agent has access to your file system. Is it safe? We don't want a background process accidentally deleting your home directory!
In the next chapter, we will learn how we lock down the specific tools this agent is allowed to use.
Next Chapter: Scoped Tool Permissions
Generated by Code IQ