Neatbo.

Text cleanup guide: encoding, lines and Unicode

Clean up text encoding, duplicates and line order, or inspect Unicode characters. Preserve paragraph and record boundaries when editing lists and documents.

Choose an operation for your task

Clean up text encoding, duplicates and line order, or inspect Unicode characters. Preserve paragraph and record boundaries when editing lists and documents.

Which operation fits?
Your taskWhere to start
Clean a list of names or IDsUse line-based deduplication or sorting after defining one record
Fix garbled textPreview a matching source encoding before exporting UTF-8
Investigate invisible differencesInspect code points before normalizing or removing characters

Work in order and retain the original

Use line editing for numbering, prefixes, lists and selected ranges. Use comparison to inspect the result, especially when whitespace or case carries meaning. Invisible-character removal can change identifiers and shaping behavior; normalization is a deliberate compatibility choice, not a universal repair button.

Representative workflow and expected result
Input: 001, 002, 001 as independent lines. Deduplication should keep the first 001 and 002; sorting is a separate decision.

Separate processing success from acceptance

Compare record counts and sample the first and last lines. Frequency counts use code points or token groups, not language-aware Chinese segmentation.

Before you finish

Expand a tool below for its steps, options and limits. Choose the tools needed for your task; you do not need to use every one.

  • Compare record counts and sample the first and last lines. Frequency counts use code points or token groups, not language-aware Chinese segmentation.
  • Reopen the download and compare it with the example and the meaning of the input.
  • When an input exceeds the stated boundaries, retain it and split the task or use a suitable processor; renaming an extension does not make it compatible.

References

Tools in this category

Expand a tool to see its steps, options and supported formats, then open its workspace.

Convert text encodingChoose the right text encoding and save a readable TXT file.

Read UTF-8, GB18030/GBK or Big5 text and export UTF-8. Inspect the decoded preview and confirm it before conversion; an incorrect but decodable legacy encoding cannot be detected reliably.

Steps

  1. Choose the local .txt file. The working copy stays in this browser.
  2. Review Input encoding, Line endings before selecting “Convert to UTF-8”.
  3. Review the UTF-8 TXT result, then use the available copy or download controls.

Available options

Input encoding
UTF-8 (strict check) · GB18030 / GBK · Big5
I checked that the decoded text is readable.
Off by default
Line endings
Keep original · LF · CRLF
Include UTF-8 BOM
Off by default

Capabilities and limits

  • Plain-text files are limited to 10 MiB each and 30 MiB total; confirm that legacy-encoding previews are readable.
Open Convert text encoding →
Merge text filesPut several text files together in the order you choose.

Join local text files in the displayed order. Choose a shared or per-file encoding, review the preview, and enter the exact separator to place between files.

Steps

  1. Choose the local .txt files in the order you want to process them. The working copy stays in this browser.
  2. Review Input encoding, Text between files before selecting “Merge text files”.
  3. Review the TXT result, then use the available copy or download controls.

Available options

Input encoding
UTF-8 (strict check) · GB18030 / GBK · Big5
I checked that the decoded text is readable.
Off by default
Text between files

Capabilities and limits

  • Plain-text files are limited to 10 MiB each and 30 MiB total; confirm that legacy-encoding previews are readable.
Open Merge text files →
Word counterWords, characters, and paragraphs at a glance.

Count pasted text by Unicode code points, UTF-16 units, grapheme clusters, Latin-letter English words, Han characters, lines, paragraphs and UTF-8 bytes. These measures differ for emoji and mixed-language text; no publisher-specific word count is implied.

Steps

  1. Paste content into “Text”, or load the example to try it.
  2. Check your input, then select “Count text”.
  3. Read the labeled counts and copy a value or the full JSON summary. Compare the required counting rule with your destination's own limit.

Capabilities and limits

  • Pasted text is limited to 1 MiB; review the result before using it.
Open Word counter →
Case converterChange letter case to suit your text.

Change pasted text with Unicode upper/lower mappings or English-oriented title, sentence and surname-first modes. The surname mode makes the first Latin word of each line uppercase and title-cases the remaining English words; review names and acronyms manually.

Steps

  1. Paste content into “Text”, or load the example to try it.
  2. Choose lowercase, uppercase, English title or sentence case, or SURNAME Given Name for one name per line; then select “Convert case”.
  3. Compare the output with the source, especially names and acronyms, then copy or download result.txt.

Available options

Case style
lowercase · UPPERCASE · Title Case (English) · Sentence case (English) · SURNAME Given Name (English)

Capabilities and limits

  • Pasted text is limited to 1 MiB; review the result before using it.
Open Case converter →
Remove duplicate linesLess repetition. Same order.

Remove repeated whole lines without changing their order. The first copy stays exactly as written; optional case and surrounding-space rules only affect comparison.

Steps

  1. Paste one entry per line into “Text”, or load the example.
  2. Choose whether comparisons ignore letter case or surrounding spaces, then remove duplicates.
  3. Check the remaining line order and copy or download the complete result.

Available options

Ignore case when comparing
Off by default
Ignore surrounding whitespace when comparing
Off by default

Capabilities and limits

  • Paste up to 1 MiB of text. This compares whole lines, not a selected column or part of a URL; review the result before using it.
  • CR, LF and CRLF line breaks use the first detected style in the result. Malformed Unicode is rejected.
Open Remove duplicate lines →
Sort text linesSort your lines as text or numbers.

Sort pasted lines as text or exact decimal values, in ascending or descending order. Equal values keep their original order; blank lines go last.

Steps

  1. Paste one item per line into “Text”, or load the numeric example.
  2. Choose Unicode lexical order or exact decimal values, then choose ascending or descending.
  3. Sort, inspect the line order, then copy or download the complete result.

Available options

Sort as
Unicode lexical order · Exact decimal values
Direction
Ascending · Descending

Capabilities and limits

  • Paste up to 1 MiB. Numeric mode requires every nonblank line to be one valid decimal number, with no labels or separators.
  • Numeric exponents must be between −1,000,000 and 1,000,000. Text order compares UTF-16 units without locale collation. Malformed Unicode is rejected.
Open Sort text lines →
Clean whitespaceTidy extra spaces and empty lines without losing useful breaks.

Clean pasted text while keeping useful line and paragraph breaks. Remove trailing spaces, optionally reduce spaces and tabs inside lines, and choose how many blank lines to retain.

Steps

  1. Paste text into “Text”, or load the example to try it.
  2. Choose trailing-space removal, space/tab collapsing, and the maximum consecutive blank lines (0 removes all).
  3. Review paragraph spacing, then copy or download the complete result.

Available options

Remove trailing spaces
On by default
Collapse spaces and tabs
Off by default
Maximum consecutive empty lines
1

Capabilities and limits

  • Paste up to 1 MiB. This tool does not join wrapped lines into paragraphs; review the spacing before using the result.
  • Trailing cleanup removes tabs and Unicode space separators. Space/tab collapsing leaves other characters intact; malformed Unicode is rejected. Mixed CR, LF and CRLF breaks use the first detected style.
Open Clean whitespace →
Find & replaceFind the text you need and replace it with care.

Replace literal matches by default. Optional JavaScript regular expressions run in a disposable worker with a hard time limit; regex replacement tokens follow JavaScript rules.

Steps

  1. Paste content into “Text”, or load the example to try it.
  2. Review Find, Replace with before selecting “Find and replace”.
  3. Review the Text result, then use the available copy or download controls.

Available options

Find
Enter as needed
Replace with
Enter as needed
Replace all matches
On by default
Match case
On by default
Use a regular expression
Off by default
Multiline anchors
Off by default

Capabilities and limits

  • Pasted text is limited to 1 MiB; review the result before using it.
Open Find & replace →
Slug generatorTurn a title into a tidy URL slug.

Normalize text with NFKC, lowercase it, and join letters and numbers with hyphens. Non-Latin letters remain by default. ASCII-only mode strips decomposable accents before removing other non-ASCII characters; it does not transliterate Chinese.

Steps

  1. Paste content into “Text”, or load the example to try it.
  2. Check your input, then select “Generate slug”.
  3. Review the URL slug result, then use the available copy or download controls.

Available options

ASCII-only slug (é becomes e; Chinese characters are removed)
Off by default

Capabilities and limits

  • Pasted text is limited to 1 MiB; review the result before using it.
Open Slug generator →
Split text fileSplit a long text file by line count or size.

Split text into line-count or UTF-8-byte-bounded parts and download a real ZIP. Byte splitting never cuts a Unicode code point; joining the parts restores the decoded source.

Steps

  1. Paste content into “TXT” or choose a local file.
  2. Review Input encoding, Split by, Lines / maximum bytes per part before selecting “Split text”.
  3. Download the ZIP, inspect the numbered TXT parts, and join their contents in order if you need to reconstruct the decoded source.

Available options

Input encoding
UTF-8 (strict check) · GB18030 / GBK · Big5
I checked that the decoded text is readable.
Off by default
Split by
Line count · UTF-8 byte budget
Lines / maximum bytes per part
1000

Switching modes resets this value to 1,000 lines or 1,048,576 UTF-8 bytes. Byte mode may split a line, but never a Unicode code point.

Capabilities and limits

  • A local TXT file is limited to 10 MiB; pasted text can be up to 30 MiB. ZIP archives contain at most 1,000 parts. Confirm that legacy-encoding previews are readable.
Open Split text file →
Email extractorExtract and deduplicate common email addresses from pasted text or a UTF-8 text file.

Find email addresses in a copied contact list, message or text/HTML source file. Review the distinct addresses, then copy or download the list.

Steps

  1. Choose Pasted text or Text file. Paste a list, message or HTML snippet, or select one UTF-8 file.
  2. Choose whether addresses that differ only in letter case should be merged; the first spelling is kept.
  3. Extract, review the count and preview, then copy or download the full one-address-per-line TXT list.

Capabilities and limits

  • Up to 1 MiB pasted text or one UTF-8 text file up to 10 MiB; up to 50,000 lines and 50,000 address matches. HTML and CSV are scanned as text, not parsed into fields.
  • Recognizes common ASCII addresses with dot-atom local parts and domain names. Quoted local parts and internationalized addresses are not covered; no mailbox, consent or deliverability check is performed.
Open Email extractor →
URL extractorPull absolute HTTP and HTTPS links from text or a UTF-8 file into a list you can copy or save.

Extract written HTTP and HTTPS links from notes, logs or source text. Review the list, then copy it or download a line-by-line TXT file. The tool does not visit the links.

Steps

  1. Choose pasted text or one UTF-8 text file, then provide content containing links.
  2. Choose HTTP, HTTPS or both, then optionally list allowed top-level domains and host domains. Decide whether to keep repeated links.
  3. Select Extract links, review the first matches, then copy or download the complete list.

Available options

Protocol
HTTP and HTTPS · HTTP only · HTTPS only
Allowed TLDs (optional)
Enter as needed

Comma-separated final labels, for example com,org. Blank allows all.

Allowed domains (optional)
Enter as needed

Comma-separated hosts, for example example.com; subdomains count too. Blank allows all.

Capabilities and limits

  • Paste up to 1 MiB or choose one UTF-8 text file up to 10 MiB. At most 50,000 lines and 50,000 matched links are processed locally.
  • Only written absolute HTTP/HTTPS links are scanned. Relative HTML hrefs and links hidden behind rich text are not resolved; no link is opened or checked for reachability.
Open URL extractor →
Line editing workspaceWrap text, number or strip lines, add affixes, create lists and extract ranges in one workspace.

Edit text line by line in one workspace. Reflow paragraphs, join soft breaks, add or remove line numbers, add prefixes or suffixes, make Markdown lists, or extract a line range. Choose one task and review the output before copying or downloading.

Steps

  1. Choose a task. For range extraction, select line numbers or literal start/end text from a static file snapshot.
  2. Paste text or select one UTF-8 text file.
  3. Run the task, inspect the result and line counts, then copy or download it.

Available options

Task
Wrap / unwrap · Add numbers · Remove numbers · Add affixes · Create list · Extract range
Line width
80
Mode
wrap · unwrap
Range selection
Line numbers · Text markers
Start at
1
Separator
.
Prefix
item-
Suffix
Enter as needed
List style
bullet · numbered · checklist
Last line (optional)
Enter as needed
Start text
Enter as needed

First line containing this exact text is included.

End text (optional)
Enter as needed

First later line containing this exact text is included; blank means through the end.

Capabilities and limits

  • Wrap joins lines within each paragraph before reflowing; blank lines separate paragraphs and long words stay intact.
  • Numbering, affixes and lists preserve a final newline without adding an extra item. Extracted ranges omit a trailing newline; blank list lines stay blank.
  • Paste up to 1 MiB or choose one UTF-8 file up to 10 MiB, with at most 50,000 physical lines. Full output is limited to 20 MiB; all option values share a 2 MiB budget. Malformed Unicode in text or active literal options is rejected.
  • Wrapping counts Unicode code points, rather than display cells or combined-character clusters. Files use strict UTF-8 and discard an initial BOM; output line breaks use LF.
Open Line editing workspace →
Word and character frequencyRank words or Unicode code points and locate a chosen word or phrase in the source text.

Paste a draft or open a UTF-8 text file to rank its words or Unicode code points. In word mode, enter a known phrase to see its count and first 20 source locations. Rankings count exact word forms; the tool does not judge writing quality.

Steps

  1. Choose Words or Code points. To locate a known phrase, enter it in word mode.
  2. Paste text or choose one UTF-8 file, then count frequencies.
  3. Inspect the ranking and phrase locations, copy matching ranking rows, or download the complete JSON/CSV ranking.

Available options

Task
Word groups · Code points
Hide common English function words
Off by default
Include spaces and line breaks
On by default
Find a word or phrase in the source (optional)
Enter as needed

Up to 120 characters. Exact case-insensitive match; the first 20 source locations are shown.

Capabilities and limits

  • Word rankings ignore case after Unicode NFC normalization. Phrase lookup is an exact case-insensitive match in the source with flexible whitespace; related forms are not grouped.
  • Character mode counts Unicode code points, including whitespace by default. One visible symbol can contain multiple code points.
  • Pasted text is limited to 1 MiB; one UTF-8 file to 10 MiB; at most 50,000 physical lines and 50,000 distinct values. All processing stays in this browser.
Open Word and character frequency →
Compare two listsFind shared entries, items only in A or B, and a combined list.

Compare two plain-text lists, one item per line. See A-only, shared and B-only groups with each preview item's first source line, so you can find it in the original list.

Steps

  1. Paste list A or choose a UTF-8 file, then paste list B. Put one item on each line.
  2. Choose the output group and decide whether case and surrounding spaces should match.
  3. Inspect all three groups and first source lines, then copy a complete group or download the selected TXT.

Available options

Download / preview
Shared · in A and B · A only · absent from B · B only · absent from A · Only in either · A-only + B-only · Combined · every distinct item
Ignore letter case
On by default
Ignore spaces around items
On by default

Capabilities and limits

  • Paste up to 1 MiB per list or choose one UTF-8 file for A (up to 10 MiB); each list is limited to 50,000 lines.
  • Blank lines and duplicates are skipped. This compares whole lines, not CSV columns, Excel workbooks or fuzzy names, and does not highlight the original sheet.
  • All processing happens in this browser; no account or upload.
Open Compare two lists →
Unicode text normalizerNormalize Unicode text to NFC, NFD, NFKC or NFKD and optionally compare two texts after normalization.

Text can look identical yet use different Unicode code points. Paste text or choose one UTF-8 file, optionally add a second text, and check whether they become identical in the selected form before copying or downloading the result.

Steps

  1. Paste text or choose one UTF-8 file. For two filenames or strings, paste the other one into the optional second input.
  2. Choose NFC or NFD for canonical equivalence; use NFKC/NFKD only when compatibility folding is intentional.
  3. Inspect the first code point change and, if supplied, the before/after comparison. Copy or download the normalized first text as UTF-8 TXT.

Available options

Normalization form
NFC · compose (general text) · NFD · decompose · NFKC · compatibility + compose · NFKD · compatibility + decompose

Capabilities and limits

  • Pasted text and the optional second text are each limited to 1 MiB; one UTF-8 file is limited to 10 MiB and 50,000 physical lines. Input stays in this browser.
  • NFKC/NFKD can replace compatibility characters and lose distinctions. Keep the original if spelling, symbols or formatting matter.
  • Normalization does not strip accents, transliterate scripts, remove invisible characters or check visually confusable text.
Open Unicode text normalizer →
Clean selected hidden charactersFind zero-width spaces and other selected hidden code points, then remove or replace only the types you choose.

Diagnose hidden characters in pasted text or a UTF-8 file. The default removes only zero-width spaces; other characters stay unless you select them. Review counts and source positions before using the result.

Steps

  1. Paste text or choose one UTF-8 text file.
  2. Choose exactly which hidden characters to remove. Optionally turn non-breaking spaces into ordinary spaces.
  3. Run the tool, inspect the per-character counts and first source positions, then copy or download the cleaned TXT.

Available options

Remove zero-width spaces · U+200B
On by default
Remove word joiners · U+2060
Off by default
Remove FEFF characters · U+FEFF
Off by default
Remove ZWNJ / ZWJ · U+200C / U+200D
Off by default
Replace non-breaking spaces with spaces · U+00A0
Off by default
Remove left-to-right marks · U+200E
Off by default

Capabilities and limits

  • Only the seven listed code points are checked; this is not a general Unicode sanitizer.
  • U+200E can affect text direction; ZWJ/ZWNJ can shape scripts and emoji. Check the original before removing them.
  • Pasted text is limited to 1 MiB; one UTF-8 text file is limited to 10 MiB and 50,000 physical lines. An initial file BOM may be consumed during UTF-8 decoding.
Open Clean selected hidden characters →
Unicode inspectorInspect each Unicode code point, its formal name or control alias, UTF-8 bytes, and source position.

Paste text or choose one UTF-8 text file to identify the exact code points behind a visible string. Search every result by code, name or bytes; export the complete inspection as JSON or CSV.

Steps

  1. Paste the text or select one UTF-8 text file.
  2. Run the inspector, then search the full result by code point, Unicode name or UTF-8 bytes. Check each source line and column.
  3. Copy matching rows as JSON or copy the full JSON, or download the complete JSON or CSV report.

Capabilities and limits

  • Up to 10,000 Unicode code points per run. Pasted text is limited to 1 MiB; one UTF-8 .txt/.md file is limited to 10 MiB and 50,000 physical lines.
  • Names come from Unicode 18.0.0. Controls use Unicode control aliases; private-use and unassigned code points have no formal name. An initial file BOM may be consumed during UTF-8 decoding.
  • Rows describe individual code points, not complete grapheme clusters or visual-confusable equivalence. Input is processed in this browser.
Open Unicode inspector →

Tools used in this article

Convert text encoding →Choose the right text encoding and save a readable TXT file.Merge text files →Put several text files together in the order you choose.Word counter →Words, characters, and paragraphs at a glance.Case converter →Change letter case to suit your text.Remove duplicate lines →Less repetition. Same order.Sort text lines →Sort your lines as text or numbers.Clean whitespace →Tidy extra spaces and empty lines without losing useful breaks.Find & replace →Find the text you need and replace it with care.Slug generator →Turn a title into a tidy URL slug.Split text file →Split a long text file by line count or size.Email extractor →Extract and deduplicate common email addresses from pasted text or a UTF-8 text file.URL extractor →Pull absolute HTTP and HTTPS links from text or a UTF-8 file into a list you can copy or save.Line editing workspace →Wrap text, number or strip lines, add affixes, create lists and extract ranges in one workspace.Word and character frequency →Rank words or Unicode code points and locate a chosen word or phrase in the source text.Compare two lists →Find shared entries, items only in A or B, and a combined list.Unicode text normalizer →Normalize Unicode text to NFC, NFD, NFKC or NFKD and optionally compare two texts after normalization.Clean selected hidden characters →Find zero-width spaces and other selected hidden code points, then remove or replace only the types you choose.Unicode inspector →Inspect each Unicode code point, its formal name or control alias, UTF-8 bytes, and source position.