Welcome to the second chapter of the FileReadTool tutorial!
In the previous chapter, Tool Definition & Interface, we defined the "Instruction Manual" that tells the AI how to ask for a file.
However, there is a danger. What if the AI asks to read a 10GB server log file?
We need a system to prevent this. We call this Resource Governance.
Think of your FileReadTool as an exclusive club.
The Bouncer stands at the door with a strict checklist. If a file is too "fat" (too many bytes) or talks too much (too many tokens), it gets turned away before it can bother the VIP.
The first step in governance is deciding what the rules are. In our project, rules aren't just hardcoded constants. We use a Hierarchy of Authority to decide the limit.
We check these sources in order:
CLAUDE_CODE_FILE_READ_MAX_OUTPUT_TOKENS?
Here is how we implement this hierarchy in limits.ts:
// File: limits.ts
export const getDefaultFileReadingLimits = memoize((): FileReadingLimits => {
// 1. Check for Feature Flags (Experiments)
const override = getFeatureValue('tengu_amber_wren', {})
// 2. Check for User Environment Variables (Highest Priority for tokens)
const envMaxTokens = getEnvMaxTokens()
// 3. calculate the final token limit
const maxTokens = envMaxTokens ?? override.maxTokens ?? 25000
return { maxTokens, /* ... other limits */ }
})
Explanation: This function is the "Brain" of the Bouncer. It gathers strictness levels from different sources and returns the final set of rules.
Let's look at a concrete example. The user asks the AI:
"Read
data_analysis.ipynb"
The file exists, but it contains thousands of generated charts and is 5MB in size. Our limit is set to 256KB.
Here is how the tool enforces this limit inside the main logic.
// File: FileReadTool.ts (Inside callInner function)
// 1. Convert notebook content to text
const cellsJson = jsonStringify(cells)
const cellsJsonBytes = Buffer.byteLength(cellsJson)
// 2. The Bouncer checks the ID card (Size in Bytes)
if (cellsJsonBytes > maxSizeBytes) {
throw new Error(
`Notebook content (${formatFileSize(cellsJsonBytes)}) exceeds ` +
`maximum allowed size (${formatFileSize(maxSizeBytes)}).`
)
}
Explanation:
By throwing an error here, we stop the process immediately. We tell the AI: "This file is too big. Try reading just a small part of it using a different tool (like jq)."
We actually have two different limits, and they protect against different things.
maxSizeBytes (The Fast Check)maxTokens (The Precise Check)Here is the implementation of the Token Guardian:
// File: FileReadTool.ts
async function validateContentTokens(content: string, ext: string, maxTokens: number) {
// 1. Make a cheap rough estimate first
const estimate = roughTokenCountEstimationForFileType(content, ext)
if (estimate <= maxTokens / 4) return // Safe!
// 2. If it looks close, do a precise count (calls an API)
const exactCount = await countTokensWithAPI(content)
if (exactCount > maxTokens) {
throw new MaxFileReadTokenExceededError(exactCount, maxTokens)
}
}
Explanation:
We don't want to run the expensive countTokensWithAPI on every tiny file. We use a heuristic (a rough guess) first. If the file seems small, we let it through. If it's borderline, we measure exactly.
When a call() is made, the tool orchestrates these checks.
call
Now let's see where this fits into the main call method we introduced in Chapter 1.
// File: FileReadTool.ts
async call(input, context) {
// 1. Load the rules
const defaults = getDefaultFileReadingLimits()
// Allow context to override defaults (e.g. for testing)
const maxSizeBytes = context.fileReadingLimits?.maxSizeBytes
?? defaults.maxSizeBytes
const maxTokens = context.fileReadingLimits?.maxTokens
?? defaults.maxTokens
// ... (Resolve file path) ...
// 2. Pass these limits down to the worker function
return await callInner(
fullFilePath,
/* ... inputs ... */,
maxSizeBytes, // <--- Passing the Bouncer's Byte rule
maxTokens, // <--- Passing the Bouncer's Token rule
/* ... other args ... */
)
}
Explanation:
The call function acts as the manager. It looks up the current laws (limits) and hands them to the worker (callInner) along with the file path. This ensures that callInner doesn't need to know how limits are calculated, only what they are.
In this chapter, we added a safety layer to our tool:
What's Next? We have secured the door. Now, once a file is allowed inside, we need to figure out what it is. Is it a text file? An image? A PDF? A Jupyter Notebook?
In the next chapter, we will build the machinery to identify and handle these different formats.
Next Chapter: Content Type Dispatcher
Generated by Code IQ