Welcome to the third chapter of the FileReadTool tutorial!
In the previous chapter, Resource Governance & Limits, we set up a "Bouncer" to stop files that are too large from crashing our system.
Now that the file has passed security and is allowed inside, we have a new problem: Not all files are the same.
We need a system to sort these files. We call this the Content Type Dispatcher.
Imagine a corporate mailroom.
The Content Type Dispatcher acts exactly like this mailroom. It looks at the file extension (the label on the package) and routes the file to the correct specialized department.
The user asks:
"Read
chart.pngandanalysis.ipynb."
Without the dispatcher, the tool would try to open the image as text, failing miserably. With the dispatcher:
.png -> Routes to the Image Processor..ipynb -> Routes to the Notebook Parser.
This logic lives inside the callInner function (which we briefly touched on in Chapter 1). This function acts as the central traffic hub.
Here is how the Dispatcher decides where to send a file request.
Let's look at how this is implemented in code. We simply check the file extension (ext) and use if/else blocks to handle the routing.
Jupyter notebooks are actually JSON files full of metadata. If we just read them as raw text, the AI gets confused by all the brackets and quotes. We parse them into a clean format first.
// File: FileReadTool.ts (Inside callInner)
if (ext === 'ipynb') {
// 1. Read and parse the complex JSON structure
const cells = await readNotebook(resolvedFilePath)
// 2. Return a specific "notebook" type object
return {
data: {
type: 'notebook',
file: { filePath: file_path, cells }
}
}
}
Explanation:
If the file ends in .ipynb, we hand it off to readNotebook. This helper function strips away the noise and gives us just the code and markdown cells the AI cares about.
If the file is an image, we can't send "text" to the AI. We need to send the image data encoded as a Base64 string (a way of representing binary data as text).
// File: FileReadTool.ts
if (IMAGE_EXTENSIONS.has(ext)) {
// 1. Read the image and maybe resize it if it's too big
const data = await readImageWithTokenBudget(resolvedFilePath, maxTokens)
// 2. Return an "image" type object
return {
data: data // contains base64 string
}
}
Explanation:
We check a list called IMAGE_EXTENSIONS (like .png, .jpg). If it matches, we call readImageWithTokenBudget. We will learn how that image processing works in Media Processing Engine.
PDFs are the most complex. Sometimes we want text, sometimes we want images of specific pages.
// File: FileReadTool.ts
if (isPDFExtension(ext)) {
// 1. Check if the user asked for specific pages (e.g., "1-5")
if (pages) {
const extractResult = await extractPDFPages(resolvedFilePath, pages)
return { data: extractResult.data }
}
// 2. Otherwise, read the whole document
const pdfData = await readPDF(resolvedFilePath)
return { data: pdfData }
}
Explanation:
The dispatcher checks for .pdf. It also checks if the user provided a pages argument (defined in our Interface in Tool Definition & Interface). It then routes to the appropriate PDF helper function.
If the file isn't a Notebook, Image, or PDF, we assume it is text (code, markdown, logs, etc.). This is the "General Delivery" of our mailroom.
// File: FileReadTool.ts
// Fallthrough: Handle as standard text
const { content, lineCount } = await readFileInRange(
resolvedFilePath,
offset, // Start line
limit // Number of lines
)
return {
data: {
type: 'text',
file: { content, numLines: lineCount, ... }
}
}
Explanation:
This acts as the else block. It calls readFileInRange, which handles reading plain text, including logic to read only specific lines (offsets) to save memory.
In Chapter 1, we defined an Output Schema (the contract). The Dispatcher ensures we fulfill that contract.
{ type: 'text', ... }.{ type: 'image', ... }.This ensures that the code receiving the result (and the AI) knows exactly what format to expect.
In this chapter, we built the Content Type Dispatcher:
callInner.What's Next? We have routed the images and PDFs to their special handlers, but how do we actually read them? How do we resize a 4K image so it doesn't cost a fortune in AI tokens?
In the next chapter, we will explore the heavy machinery behind handling binary files.
Next Chapter: Media Processing Engine
Generated by Code IQ