In Chapter 1: Telemetry Bootstrap & Instrumentation, we set up our "recording studio"βthe Providers and Exporters. Now that the equipment is ready, we need to record the actual "music" of our application.
In an asynchronous application like Claude Code (CLI), many things happen at once. A user types a prompt, the AI thinks, a tool executes, and files are read. Session Tracing is the art of connecting all these separate actions into a single, coherent story.
To understand tracing, think of Russian Nesting Dolls.
curl weather.api).If the "Inner Doll" (Tool) breaks, we need to know exactly which "Outer Doll" (User Interaction) it belonged to. This linking process is called Context Propagation.
Imagine a user runs this command:
"Create a new file called hello.txt"
We want our telemetry to generate a timeline that looks like this:
[Interaction Span: "Create hello.txt"] (Total: 2000ms)
βββ [LLM Request Span] (Thinking... 500ms)
βββ [Tool Span: FileCreator] (Writing... 100ms)
βββ [LLM Request Span] (Confirming... 200ms)
Without tracing, these would just be four unconnected log messages. With tracing, they are a family tree.
A Span is a single operation with a start time and an end time. It represents one "doll."
This is the invisible thread that connects a child to its parent. When code executes await doSomething(), the system must remember "Who called me?" so the child span creates the correct link.
In Node.js/TypeScript, we use a tool called AsyncLocalStorage. Think of it as a backpack that automatically follows your code execution. If you put the "Parent Span" in the backpack at the start, any function called later can look inside the backpack to find its parent.
The sessionTracing.ts module provides helper functions to manage these dolls easily.
When the user hits "Enter", we start the main span.
import { startInteractionSpan, endInteractionSpan } from './sessionTracing'
// 1. Start the stopwatch for the user's request
const span = startInteractionSpan("Create a file")
try {
// ... Application logic runs here ...
} finally {
// 2. Stop the stopwatch when done
endInteractionSpan()
}
Inside the interaction, we might call the LLM. Notice we don't manually pass the parent! The "backpack" handles it.
import { startLLMRequestSpan, endLLMRequestSpan } from './sessionTracing'
// The system automatically knows this is inside the Interaction
const llmSpan = startLLMRequestSpan("claude-3-5-sonnet")
// Call the AI API...
await callAnthropicAPI()
// Stop the stopwatch, recording success or failure
endLLMRequestSpan(llmSpan, {
inputTokens: 50,
success: true
})
If the LLM decides to run a tool, we wrap that too.
import { startToolSpan, endToolSpan } from './sessionTracing'
// Start the tool timer
startToolSpan("file_writer", { filename: "hello.txt" })
// Do the actual file writing...
await fs.writeFile("hello.txt", "Hello!")
// Stop the tool timer
endToolSpan("Success")
What happens under the hood when we call startInteractionSpan?
AsyncLocalStorage (the backpack).
Let's look at sessionTracing.ts to see how this is built.
We create a storage container. This is a built-in Node.js feature that is "async-aware."
// From sessionTracing.ts
import { AsyncLocalStorage } from 'async_hooks'
// Holds the current active span
const interactionContext = new AsyncLocalStorage<SpanContext | undefined>()
When we start a span, we enter the context.
export function startInteractionSpan(userPrompt: string): Span {
const tracer = getTracer()
// 1. Create the span using OpenTelemetry
const span = tracer.startSpan('claude_code.interaction', {
attributes: { user_prompt: userPrompt }
})
// 2. Prepare the object to store
const spanContextObj = { span, startTime: Date.now(), attributes: {} }
// 3. Put it in the backpack!
// All code following this line will see this span as "active"
interactionContext.enterWith(spanContextObj)
return span
}
Beginner Note:
interactionContext.enterWith(...)is the critical line. It says, "For the rest of this asynchronous execution chain, this object is our global state."
When a child (like a tool) starts, it checks the backpack.
export function startToolSpan(toolName: string): Span {
const tracer = getTracer()
// 1. Look inside the backpack
const parentSpanCtx = interactionContext.getStore()
// 2. If a parent exists, link them!
const ctx = parentSpanCtx
? trace.setSpan(otelContext.active(), parentSpanCtx.span)
: otelContext.active()
// 3. Start the child span associated with that context
return tracer.startSpan('claude_code.tool', { attributes: { tool_name: toolName } }, ctx)
}
You might notice maps like activeSpans using WeakRef in the full code.
const activeSpans = new Map<string, WeakRef<SpanContext>>()
This is a safety mechanism. If a developer forgets to call endInteractionSpan(), the WeakRef allows the memory to be freed (Garbage Collected) automatically, preventing memory leaks in long-running applications.
Sometimes standard tracing isn't enough. We might want to see the exact system prompt used or the full output of the LLM. This is handled in betaSessionTracing.ts.
It uses a clever Hashing Strategy to avoid sending too much data.
// From betaSessionTracing.ts
// Instead of sending a 50kb system prompt every time...
const promptHash = hashSystemPrompt(newContext.systemPrompt)
// We check if we've seen it before
if (!seenHashes.has(promptHash)) {
// If new, send the full prompt
logOTelEvent('system_prompt', { text: fullPrompt })
seenHashes.add(promptHash)
} else {
// If seen, just send the hash ID (saves bandwidth!)
span.setAttribute('system_prompt_hash', promptHash)
}
In this chapter, we learned:
Now that we have the timeline (Spans) set up, we often need to record specific point-in-time occurrences, like "User clicked a button" or "Error occurred." For that, we need Logs.
Next Chapter: Discrete Event Logging
Generated by Code IQ