Neatbo.

Unicode text normalizer

Normalize Unicode text to NFC, NFD, NFKC or NFKD and optionally compare two texts after normalization.

Browser-local processingInputTXTOutputTXTUp to 10 MiB per file · File limit: 1
  1. 1Add input
  2. 2Adjust settings
  3. 3Get your result

Tool input and files are processed in this browser without being uploaded.

Your input

Inputs are kept temporarily in this tab when switching tools. Refreshing or closing clears them; large results may need to be regenerated.

⌘ / Ctrl + Enter to run

Choose normalization

Pick the target form before processing. NFC is the general-text default.

NFC composes and NFD decomposes canonically equivalent sequences. Characters may look unchanged while their code points differ.

Normalize from
0 characters · 0 bytes
Preparing the tool…

Before you start

Text can look identical yet use different Unicode code points. Paste text or choose one UTF-8 file, optionally add a second text, and check whether they become identical in the selected form before copying or downloading the result.

How to use this tool

  1. Paste text or choose one UTF-8 file. For two filenames or strings, paste the other one into the optional second input.
  2. Choose NFC or NFD for canonical equivalence; use NFKC/NFKD only when compatibility folding is intentional.
  3. Inspect the first code point change and, if supplied, the before/after comparison. Copy or download the normalized first text as UTF-8 TXT.

Supported inputs and limits

Pasted text and the optional second text are each limited to 1 MiB; one UTF-8 file is limited to 10 MiB and 50,000 physical lines. Input stays in this browser.

NFKC/NFKD can replace compatibility characters and lose distinctions. Keep the original if spelling, symbols or formatting matter.

Normalization does not strip accents, transliterate scripts, remove invisible characters or check visually confusable text.

Worked example

Example input

Café
Example options
{"form":"NFC"}

Example output

Café

When something does not work

Pasted text is limited to 1 MiB; one UTF-8 text file is limited to 10 MiB and 50,000 physical lines. Input stays in this browser.

Frequently asked questions

Why do the characters still look the same?

Canonical normalization can change code points without changing appearance. The result shows input/output code point counts and the first changed sequence. If it says unchanged, the text was already in the selected form.

Which form should I choose?

NFC is the general-text default. NFD separates composed characters. NFKC and NFKD also fold compatibility characters, such as circled digits or ligatures, and can lose meaningful distinctions.

Does this remove accents or make names safe to compare?

You can compare two manually entered strings after normalization. This does not remove accents, fold case, transliterate, detect visual confusables or batch-rename files.

Is my input uploaded?

No. The text or chosen file is processed in this browser. Save the result before refreshing or closing the tab.

Documentation & further reading

Related tools