In the previous chapter, Session State Management, we gave our CLI a memory. We learned how to save the "tag" to a JSON file so the computer remembers it later.
But here is the catch: Computers are very literal.
If a user copies and pastes a word from a website, they might accidentally paste invisible characters (like a "Zero Width Space"). To a human, tag and tag look the same. To a computer, one might be tag and the other t\u200bag. If we save the "dirty" version, our search feature will fail later because the keys don't match.
In this chapter, we explore Input Sanitization.
Imagine your application's database is a pristine water tank. The user input is water coming from a river. Usually, it's clean, but sometimes it contains invisible bacteria or dirt.
Input Sanitization is the filter system.
In our tag.tsx file, we don't just take args (the user input) and save it immediately. We pass it through a cleaning function first.
We use a utility function called recursivelySanitizeUnicode.
// tag.tsx
import { recursivelySanitizeUnicode } from '../../utils/sanitization.js';
Inside our component, before we do anything else (before checking if the tag exists, or asking for confirmation), we clean the input.
function ToggleTagAndClose({ tagName, onDone }) {
// 1. Sanitize: Remove invisible characters
// 2. Trim: Remove spaces from start/end
const normalizedTag = recursivelySanitizeUnicode(tagName).trim();
// ... rest of the logic uses 'normalizedTag', NOT 'tagName'
Explanation:
" bugfix "."bugfix".
Now, whenever we save to the database or compare strings, we use normalizedTag.
What exactly does recursivelySanitizeUnicode do? It creates a safe, standard version of text.
รฉ which can be written as e + ') into a single standard character.
Let's look at a simplified version of the internal code in utils/sanitization.js.
First, it handles Normalization. Unicode allows multiple ways to write the same character. We want the standard form ("NFC").
// utils/sanitization.js (Simplified)
const sanitizeString = (str: string): string => {
// 1. Convert to Standard Form (NFC)
// This turns decomposed characters into single ones.
return str.normalize('NFC')
// 2. Remove "Control Characters" (Invisible codes 0-31)
.replace(/[\u0000-\u001f]/g, '');
};
However, our data might not just be a string. It might be an object, like { id: 1, name: "tag" }. That is why the function is Recursive. It digs into objects to find strings.
export function recursivelySanitizeUnicode(value: unknown): any {
// Case A: It's a string -> Clean it
if (typeof value === 'string') {
return sanitizeString(value);
}
// Case B: It's an array -> Clean every item
if (Array.isArray(value)) {
return value.map(recursivelySanitizeUnicode);
}
// Case C: It's an object -> Clean every value
if (value && typeof value === 'object') {
// ... logic to loop through object keys
}
}
Explanation:
In this chapter, we learned:
recursivelySanitizeUnicode to strip out bad characters and normalize text.tag === tag is always true.Now that our application is registered, has a UI, saves data, and cleans that data, we need to know if people are actually using it!
Generated by Code IQ