Unicode inspector
Inspect each Unicode code point, its formal name or control alias, UTF-8 bytes, and source position.
- 1Add input
- 2Review and run
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Paste text or choose one UTF-8 text file to identify the exact code points behind a visible string. Search every result by code, name or bytes; export the complete inspection as JSON or CSV.
How to use this tool
- Paste the text or select one UTF-8 text file.
- Run the inspector, then search the full result by code point, Unicode name or UTF-8 bytes. Check each source line and column.
- Copy matching rows as JSON or copy the full JSON, or download the complete JSON or CSV report.
Supported inputs and limits
Up to 10,000 Unicode code points per run. Pasted text is limited to 1 MiB; one UTF-8 .txt/.md file is limited to 10 MiB and 50,000 physical lines.
Names come from Unicode 18.0.0. Controls use Unicode control aliases; private-use and unassigned code points have no formal name. An initial file BOM may be consumed during UTF-8 decoding.
Rows describe individual code points, not complete grapheme clusters or visual-confusable equivalence. Input is processed in this browser.
Worked example
Example input
A中
Example options
{}Example output
[
{
"index": 1,
"line": 1,
"column": 1,
"character": "A",
"codepoint": "U+0041",
"name": "LATIN CAPITAL LETTER A",
"nameType": "formal",
"utf8": "41"
},
{
"index": 2,
"line": 1,
"column": 2,
"character": "中",
"codepoint": "U+4E2D",
"name": "CJK UNIFIED IDEOGRAPH-4E2D",
"nameType": "formal",
"utf8": "e4 b8 ad"
}
]When something does not work
If a text file fails to open, confirm it is valid UTF-8. If the input exceeds 10,000 code points, inspect a smaller excerpt. Keep the original for comparison.
Frequently asked questions
Why can one visible symbol have several rows?
Emoji sequences and combining marks may contain multiple Unicode code points. Each row represents one code point in source order; the visual placeholder is not a replacement character.
Are names available for every code point?
Unicode 18.0.0 formal names are shown for assigned characters. Controls use official control aliases. Private-use and unassigned code points have no formal name.
Does this compare two strings or detect lookalike attacks?
No. It lists one input’s code points, names, bytes and positions. Compare two reports yourself; visual-confusable detection and normalization are separate tasks.
Is my input sent to a server?
No. Text and selected files are processed in your browser. Save the report you need before refreshing or closing the tab.
Documentation & further reading
Related tools
Convert text encoding
Choose the right text encoding and save a readable TXT file.
Merge text files
Put several text files together in the order you choose.
Word counter
Words, characters, and paragraphs at a glance.
Case converter
Change letter case to suit your text.