Welcome back! In Chapter 2: The Executor (Computer Control), we gave the AI the power to control your mouse and keyboard. It can now click buttons, type text, and drag files.
But as Spider-Man knows: "With great power comes great responsibility."
What happens if the AI gets confused and starts clicking wildly? Or gets stuck in a loop opening 100 Calculator windows? You need a way to stop it immediately.
The Safety & Abort Mechanism is a global "Kill Switch."
Imagine the AI is trying to drag a file, but it miscalculates and keeps holding the mouse button down while moving the cursor endlessly.
The Goal: You, the user, press the Escape key on your physical keyboard. The system immediately detects thisβno matter which window is focusedβand completely shuts down the AI's current action.
To build this safety net, we need to understand three concepts:
Escape even if you are using Chrome or Spotify.Escape key before the current app sees it. This prevents the AI from accidentally dismissing a dialog box when you actually intended to stop the AI.Using this mechanism is designed to be simple. We register a "callback" function that runs whenever the user hits Escape.
We usually set this up when the session starts.
// From escHotkey.ts
import { registerEscHotkey } from './escHotkey'
// Start listening for the Escape key
const success = registerEscHotkey(() => {
console.log("ABORT! User pressed Escape!")
// Logic to stop the executor goes here
})
Explanation: We pass a function to registerEscHotkey. If the user presses Escape, that function runs immediately.
When the AI finishes its task successfully, we stop listening so the Escape key goes back to normal behavior.
// From escHotkey.ts
import { unregisterEscHotkey } from './escHotkey'
// Stop listening
unregisterEscHotkey()
Explanation: Always clean up after yourself! If we don't unregister, the user won't be able to use their Escape key normally.
How does the system know the difference between you pressing Escape and the AI pressing Escape?
We use a technique I call the "I'm Doing It" Notification. Before the AI presses Escape, it raises its hand and tells the safety system, "Hey, this next key press is me. Don't abort."
Let's look at escHotkey.ts to see how this is implemented.
We use a native Swift module to tap into the OS event stream.
// From escHotkey.ts
let registered = false
export function registerEscHotkey(onEscape: () => void): boolean {
if (registered) return true
// Call the native Swift code to listen for KeyDown
const cu = requireComputerUseSwift()
if (!cu.hotkey.registerEscape(onEscape)) {
return false // Failed (maybe permissions issue)
}
registered = true
return true
}
Explanation: cu.hotkey.registerEscape is the bridge to the Operating System. It tells macOS: "Send me a signal whenever Escape is pressed."
This is the clever part. We export a function called notifyExpectedEscape.
// From escHotkey.ts
export function notifyExpectedEscape(): void {
if (!registered) return
// Tell native code: "The next Escape is from the AI"
requireComputerUseSwift().hotkey.notifyExpectedEscape()
}
Explanation: This sets a temporary flag (usually with a tiny timeout, like 100ms). If an Escape event comes in while that flag is up, it's ignored.
Now, let's look back at Chapter 2's executor.ts. When the AI wants to type a key, it checks if that key is "Escape".
// From executor.ts
async key(keySequence: string): Promise<void> {
const parts = keySequence.split('+')
// Check: Is the AI trying to press Escape?
const isEsc = isBareEscape(parts)
await drainRunLoop(async () => {
// If yes, warn the safety system first!
if (isEsc) {
notifyExpectedEscape()
}
// Now actually press the key
await input.keys(parts)
})
}
Explanation:
Escape.notifyExpectedEscape() defined in our safety module.In this chapter, we added a critical layer of safety to our application:
Now the AI can control the computer (Chapter 2) and we can stop it if it goes rogue (Chapter 3). But how do these pieces effectively talk to the Operating System to find windows and take screenshots?
Generated by Code IQ