In the previous chapter, Perfetto Performance Profiling, we acted like mechanics in the engine room, using high-precision tools to debug performance issues locally on a single machine.
But what if the "Captain" (Product Management) asks: "Are users running into more errors today than yesterday across the entire world?"
We cannot look at every user's local trace file. We need to ship summary data from thousands of computers to a central warehouse where we can run SQL queries.
Welcome to the Custom Metrics Export Strategy.
Think of OpenTelemetry (OTel) as a factory that boxes up data (metrics) into standard shipping containers.
However, our destinationβGoogle BigQueryβis very picky. It doesn't have a dock for standard containers. It requires:
The BigQueryMetricsExporter is our Specialized Courier. It takes the standard boxes, unpacks them, fills out the BigQuery paperwork, repackages them, and drives them to the warehouse.
Imagine we want to build a dashboard that shows:
"Total tokens generated by Claude Code users per hour."
To make this happen:
tokens_generated.add(50).This is the job title. In OpenTelemetry, an "Exporter" is a class responsible for sending data away. "Push" means we actively send data on a schedule (e.g., every 60 seconds), rather than waiting for the server to ask for it.
This is the most important technical concept in this chapter.
BigQuery likes Delta. If we send "10,000" and then "10,005", BigQuery might add them together (20,005), which is wrong. By sending Delta, we simply send "5", and BigQuery adds that to the total.
The raw OTel data is a complex tree of objects. BigQuery wants a flat list. The exporter acts as a "flattening" machine.
Before looking at the code, let's visualize the journey of a metric.
We implement this in bigqueryExporter.ts. Let's break it down into small, digestible steps.
Before we act as a courier, we check if we are allowed to deliver the package. We check if the user has trusted the app and if they have opted out of telemetry.
// From bigqueryExporter.ts
private async doExport(metrics: ResourceMetrics, callback: any): Promise<void> {
// 1. Check if the user has accepted the trust dialog
const hasTrust = checkHasTrustDialogAccepted()
// 2. Check if the organization has disabled metrics
const metricsStatus = await checkMetricsEnabled()
if (!hasTrust || !metricsStatus.enabled) {
// If not allowed, pretend we succeeded but do nothing
callback({ code: ExportResultCode.SUCCESS })
return
}
// ... continue to export ...
}
This is the heart of the "Courier" work. We take the raw metrics object and reshape it into InternalMetricsPayload.
We extract the Resource Attributes (Who is this?) and the Data Points (What happened?).
// Inside transformMetricsForInternal()
const resourceAttributes = {
// Identify the sender
'service.name': 'claude-code',
'os.type': attrs['os.type'] || 'unknown',
// Identify the customer type (Free vs Paid)
'user.customer_type': isClaudeAISubscriber() ? 'claude_ai' : 'api'
}
// Map the complex OTel structure to simple JSON
const payload = {
resource_attributes: resourceAttributes,
metrics: metrics.scopeMetrics.flatMap(scope =>
scope.metrics.map(metric => ({
name: metric.descriptor.name,
data_points: this.extractDataPoints(metric)
}))
)
}
A metric might have many data points. We filter out anything that isn't a number and convert the timestamp to a human-readable string.
private extractDataPoints(metric: MetricData): DataPoint[] {
return metric.dataPoints
// Ensure we only send numbers
.filter((p): p is OTelDataPoint<number> => typeof p.value === 'number')
// Create the clean data object
.map(point => ({
attributes: this.convertAttributes(point.attributes),
value: point.value,
timestamp: new Date().toISOString() // Simplified for tutorial
}))
}
Finally, we put the JSON payload in an HTTP POST request and send it to the Anthropic API (which forwards it to BigQuery).
// Inside doExport()
// Get security headers (Auth tokens)
const authResult = getAuthHeaders()
const response = await axios.post(
this.endpoint,
payload,
{
headers: {
'Content-Type': 'application/json',
...authResult.headers, // Attach the ID badge
}
}
)
As mentioned in the concepts, sending "Total since startup" (Cumulative) messes up our database math. We must force the system to reset the counter to 0 after every export.
We do this by overriding a specific configuration method in the class.
selectAggregationTemporality(): AggregationTemporality {
// DO NOT CHANGE THIS TO CUMULATIVE
// If we do, the dashboard graphs will show massive, incorrect spikes.
return AggregationTemporality.DELTA
}
Beginner Note:
DELTAmeans if you count 5 errors, send "5". The next time you count 3 errors, send "3".
CUMULATIVEwould send "5", and then the next time send "8" (5+3).
Congratulations! You have completed the Telemetry tutorial series. Let's look at the full picture of what you have built:
You now understand the complete lifecycle of observability in a modern production application. From the moment a user presses a key, to the moment a graph updates on an executive dashboard, you own the pipeline.
Generated by Code IQ