In the previous chapter, Conversation Summarization (Compaction), we learned how to take a long conversation and rewrite it into a concise summary.
But here is the catch: You cannot just cut a conversation anywhere you want.
If you cut the history at the wrong spotβlike in the middle of a sentence or halfway through a complex taskβyou will confuse the AI and cause API errors. In this chapter, we will look at Message Grouping, the utility that ensures we only slice the conversation at "safe" points.
Imagine you are on a phone call.
Now, imagine the Auto-Compact system wakes up at step 3 because the memory is full. It decides to summarize everything up to that point and delete the recent messages.
If it cuts the history between step 2 and step 4, the AI wakes up with amnesia. It sees the "Tool Output: 2500" but has zero memory of asking for it. This causes a dangling_tool_result error, and the API crashes.
Use Case: You are building an autonomous agent that performs multi-step coding tasks. It edits a file, runs a test, sees the error, fixes the file, and runs the test again. You need to summarize the history, but you must ensure you don't delete the "Run Test" command while keeping the "Test Result."
To prevent these crashes, we group messages into logical units called API Rounds.
We use a function called groupMessagesByApiRound. It takes a flat list of messages and returns a list of "groups."
1. User: "Check the weather."
2. Assistant (ID: A1): "Using weather tool..."
3. Tool (ID: A1): "Sunny"
4. Assistant (ID: A1): "It is sunny."
5. User: "Great, thanks."
6. Assistant (ID: B2): "You're welcome."
The system groups these based on the interaction flow.
const groups = groupMessagesByApiRound(allMessages);
The function returns an array of arrays. Notice how the Tool Call and Result stay together with the Assistant that requested them.
Group 1:
[User: "Check weather", Assistant: "Using tool...", Tool: "Sunny", Assistant: "It is sunny"]
Group 2:
[User: "Great, thanks", Assistant: "You're welcome"]
How does the system know where to draw the line? It watches the Assistant ID.
When the AI replies, it generates a unique ID. If it calls a tool and then continues speaking, that ID usually remains consistent (or logically linked) for that "turn." If we see a new Assistant ID, we know a new round has started.
grouping.tsLet's look at the actual code that handles this. It is a simple loop with a "state" tracker.
First, we set up our containers.
export function groupMessagesByApiRound(messages: Message[]): Message[][] {
const groups: Message[][] = []
let currentGroup: Message[] = []
// We track the ID of the last assistant we saw
let lastAssistantId: string | undefined
// ... loop starts here
}
Explanation: groups will hold our final result. currentGroup builds the specific chain we are currently looking at.
Next, we iterate through every message. This is where the magic happens. We check if the Assistant ID has changed.
for (const msg of messages) {
// Check if this is a NEW assistant turn
const isNewTurn = msg.type === 'assistant' &&
msg.message.id !== lastAssistantId &&
currentGroup.length > 0
if (isNewTurn) {
// Seal the previous group and start a new one
groups.push(currentGroup)
currentGroup = [msg]
} else {
// Otherwise, keep adding to the current chain
currentGroup.push(msg)
}
Explanation: If we see an Assistant message, and its ID is different from the last one we saw, we know the previous conversation "round" is finished. We save that group and start fresh.
Finally, we update our tracker and handle the end of the loop.
// Update the tracker so we know for next time
if (msg.type === 'assistant') {
lastAssistantId = msg.message.id
}
}
// Don't forget the very last group!
if (currentGroup.length > 0) {
groups.push(currentGroup)
}
return groups
}
Explanation: We always update lastAssistantId so we can compare it against the next message. After the loop finishes, we push whatever is left in currentGroup to the final list.
Now that we have these groups, the Compaction engine (Chapter 3) becomes much safer.
Instead of saying "Delete messages 1 through 15," the engine can say "Delete Groups 1 through 3."
This ensures that we never accidentally delete a Tool Call while leaving the Tool Result dangling. We either keep the whole interaction, or we summarize the whole interaction.
In this chapter, you learned:
Now we can group messages safely. But sometimes, even a single group is too large. What if the AI generated 500 lines of JSON data in one turn? We don't want to summarize the whole conversation, we just want to shrink that one massive message.
Next Chapter: Micro-Compaction & Pruning
Generated by Code IQ