Welcome back! In the previous chapter, Post-Sampling Extraction Hook, we created the "Court Stenographer" (the hook) that wakes up after every message to see if it needs to work.
However, we left one major question unanswered: How does the Stenographer decide when to type?
If the AI summarizes the conversation after every single message ("User said hi", "I said hi back"), it wastes electricity, costs money, and creates a messy, repetitive memory file. We need a set of rulesβa Threshold Logicβto determine when enough "significant" things have happened to justify an update.
In this chapter, we will build the brain of our Stenographer.
Imagine a user is coding with the AI.
To make this decision, we look at three specific metrics (counters).
A "token" is roughly part of a word. LLMs charge by the token. We track how much the conversation has grown.
When an AI uses a tool (like reading a file or running a terminal command), it usually means it's doing "real work."
If the AI is in the middle of a complex chain of thoughts (using a tool right now), we shouldn't interrupt it to write notes. We prefer to update when the AI has finished a thought and is waiting for the user.
Here is the decision tree our logic follows.
The Golden Rule: We always require the token threshold (growth) to be met. Once that is met, we look for either high tool usage OR a quiet moment to perform the update.
The logic lives in sessionMemory.ts and sessionMemoryUtils.ts. Let's break down the shouldExtractMemory function.
When a session starts, we don't want to create a memory file for a conversation that might only be 2 messages long. We wait for a "bulk" of context (e.g., 10,000 tokens) before we create the first entry.
// Inside shouldExtractMemory()
const currentTokenCount = tokenCountWithEstimation(messages)
// If we haven't started memory yet...
if (!isSessionMemoryInitialized()) {
// Check if we hit the "Big Bang" threshold (e.g., 10k tokens)
if (!hasMetInitializationThreshold(currentTokenCount)) {
return false // Not yet!
}
markSessionMemoryInitialized()
}
Once initialized, we switch to "Incremental Mode." We track how much the conversation has grown since the last time we updated.
// In sessionMemoryUtils.ts
export function hasMetUpdateThreshold(currentTokenCount: number): boolean {
// Calculate growth: Current Size - Size At Last Update
const tokensSinceLastExtraction =
currentTokenCount - tokensAtLastExtraction
// e.g., Return true if we grew by 5,000 tokens
return tokensSinceLastExtraction >= config.minimumTokensBetweenUpdate
}
tokensAtLastExtraction every time we write to the file. This function simply checks the difference.
We also count how many times the AI used a tool (like FileRead or RunCommand) since the last update.
// Count tools used since the last specific message ID we remembered
const toolCallsSinceLastUpdate = countToolCallsSince(
messages,
lastMemoryMessageUuid,
)
// Check against our config (e.g., 3 tools)
const hasMetToolCallThreshold =
toolCallsSinceLastUpdate >= getToolCallsBetweenUpdates()
Now we combine these factors. This is the most critical logic in the system.
// Check if the AI just finished using a tool in the very last message
const hasToolCallsInLastTurn = hasToolCallsInLastAssistantTurn(messages)
// THE DECISION LOGIC
const shouldExtract =
// Case A: Lots of tokens AND lots of tool usage (Update immediately!)
(hasMetTokenThreshold && hasMetToolCallThreshold) ||
// Case B: Lots of tokens AND it's a "quiet" moment (Good time to write)
(hasMetTokenThreshold && !hasToolCallsInLastTurn)
Let's analyze this logic:
hasMetTokenThreshold is mandatory. If the conversation hasn't grown enough (e.g., only 100 new tokens), we never update, regardless of tool usage. This prevents spamming updates.
If shouldExtract is true, we need to prepare for the next cycle.
if (shouldExtract) {
// Save the ID of the last message so we know where to start counting next time
const lastMessage = messages[messages.length - 1]
if (lastMessage?.uuid) {
lastMemoryMessageUuid = lastMessage.uuid
}
return true
}
return false
You have successfully implemented the Update Threshold Logic.
Our system is now efficient.
Now the "Court Stenographer" knows when to type. But waitβtyping a summary is hard work! If the main AI stops to read the whole history and summarize it, the user will see a loading spinner. That's bad user experience.
We need a way to do this work in the background without blocking the user.
In the next chapter, we will learn how to spawn a "Clone" of our AI to do the heavy lifting.
Next Chapter: Isolated Forked Agent
Generated by Code IQ