In the previous chapter, Environment Layout Management, we organized our agents into neat, colorful windows.
However, a pretty window doesn't make an agent safe. Imagine you spawn a "Cleanup Agent" and it decides to run rm -rf / (delete everything) on your computer. If the agent is running in a background process or a separate terminal window, how do you stop it?
We need a way for the Worker (the agent) to ask the Leader (you/the main app) for permission before doing something dangerous. This is the Distributed Permission System.
Think of your application like a corporate office.
If the Junior Employee sits in a different building (a separate process like tmux), they can't just tap you on the shoulder. They need a formal process:
The Distributed Permission System digitizes this workflow using JSON files and file watching.
This is a data object that describes intent. It contains:
ls -la /private/logsSince the Leader and the Worker might be in completely different memory spaces (or even different machines in the future), they communicate via the filesystem. The Worker writes a request file to a specific folder, and the Leader watches that folder.
When an agent asks for permission, it effectively "freezes." It enters a loop where it checks the mailbox every 500ms: "Did the boss reply yet? ... Did the boss reply yet?"
As a developer using Swarm, you rarely write the polling logic yourself. It is built into the canUseTool hook provided to the agents.
Here is what the flow looks like from the Worker's perspective.
The agent tries to use a tool, like BashTool. The system interrupts before the tool executes.
// Inside the agent runner logic
const result = await canUseTool(
toolName,
inputArguments
);
// The code PAUSES here until the user approves in the UI!
if (result.behavior === 'allow') {
executeTool();
}
What happens here: The canUseTool function handles the entire complex dance of creating a JSON file, waiting for the user to click "Approve" in the UI, and reading the response.
Let's visualize how a "Remote" agent gets permission from the "Local" user.
Let's explore the code files that power this system.
permissionSync.ts)When a worker needs permission, it creates a standardized object. This ensures the Leader UI knows exactly what to display.
// permissionSync.ts
export function createPermissionRequest(params) {
return {
id: generateRequestId(), // e.g., "perm-12345"
workerName: params.workerName,
toolName: params.toolName, // e.g., "Bash"
input: params.input, // e.g., { command: "ls" }
status: 'pending',
createdAt: Date.now(),
}
}
permissionSync.ts)Instead of direct function calls, we serialize this object to JSON and write it to the Leader's mailbox file.
// permissionSync.ts
export async function sendPermissionRequestViaMailbox(request) {
const leaderName = await getLeaderName(request.teamName);
// Write the JSON message to the leader's folder
await writeToMailbox(leaderName, {
from: request.workerName,
text: JSON.stringify(request),
timestamp: new Date().toISOString()
});
}
inProcessRunner.ts)The worker cannot proceed until it gets an answer. It enters a polling loop.
// inside createInProcessCanUseTool
const pollInterval = setInterval(async () => {
// 1. Check mailbox for new messages
const messages = await readMailbox(identity.agentName);
// 2. Look for a response to OUR specific request ID
const response = findResponse(messages, requestId);
if (response) {
clearInterval(pollInterval);
resolve(response.decision); // 'approved' or 'rejected'
}
}, 500); // Check every half second
Beginner Note: This pattern is often called "Long Polling." It's simple and robust. If the Leader crashes and restarts, the file is still there, so the request isn't lost!
leaderPermissionBridge.ts)If the agent happens to be running In-Process (inside the same app), we can sometimes skip the file system for speed. The "Bridge" connects the backend logic directly to the React Frontend.
// leaderPermissionBridge.ts
// A global variable stores the function to update the React UI
let registeredSetter = null;
export function registerLeaderToolUseConfirmQueue(setter) {
// The React component calls this to say "I am ready to show popups"
registeredSetter = setter;
}
This acts like a direct phone line. If the phone line is connected (In-Process), the agent calls directly. If not (Remote/Tmux), the agent sends a letter (Mailbox).
The Distributed Permission System ensures security across boundaries.
This system makes your AI swarm safe to use. You can have 5 agents working in parallel, and if any of them tries to do something risky, they will politely wait for your go-ahead.
Now that our agents are running, organized, and secure, there is one big question left: What happens if I turn off my computer? Does the team forget everything?
In the next chapter, we will learn how we save the brain of the operation.
Next Chapter: Team State Persistence
Generated by Code IQ