Text cleanup guide: encoding, lines and Unicode
Clean up text encoding, duplicates and line order, or inspect Unicode characters. Preserve paragraph and record boundaries when editing lists and documents.
Choose an operation for your task
Clean up text encoding, duplicates and line order, or inspect Unicode characters. Preserve paragraph and record boundaries when editing lists and documents.
| Your task | Where to start |
|---|---|
| Clean a list of names or IDs | Use line-based deduplication or sorting after defining one record |
| Fix garbled text | Preview a matching source encoding before exporting UTF-8 |
| Investigate invisible differences | Inspect code points before normalizing or removing characters |
Work in order and retain the original
Use line editing for numbering, prefixes, lists and selected ranges. Use comparison to inspect the result, especially when whitespace or case carries meaning. Invisible-character removal can change identifiers and shaping behavior; normalization is a deliberate compatibility choice, not a universal repair button.
Input: 001, 002, 001 as independent lines. Deduplication should keep the first 001 and 002; sorting is a separate decision.Separate processing success from acceptance
Compare record counts and sample the first and last lines. Frequency counts use code points or token groups, not language-aware Chinese segmentation.
Before you finish
Expand a tool below for its steps, options and limits. Choose the tools needed for your task; you do not need to use every one.
- Compare record counts and sample the first and last lines. Frequency counts use code points or token groups, not language-aware Chinese segmentation.
- Reopen the download and compare it with the example and the meaning of the input.
- When an input exceeds the stated boundaries, retain it and split the task or use a suitable processor; renaming an extension does not make it compatible.
References
- Text — format and implementation reference
Reference for format and processing boundaries. The operating reference states the actual supported scope of this site.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Convert text encodingChoose the right text encoding and save a readable TXT file.
Read UTF-8, GB18030/GBK or Big5 text and export UTF-8. Inspect the decoded preview and confirm it before conversion; an incorrect but decodable legacy encoding cannot be detected reliably.
Steps
- Choose the local .txt file. The working copy stays in this browser.
- Review Input encoding, Line endings before selecting “Convert to UTF-8”.
- Review the UTF-8 TXT result, then use the available copy or download controls.
Available options
- Input encoding
- UTF-8 (strict check) · GB18030 / GBK · Big5
- I checked that the decoded text is readable.
- Off by default
- Line endings
- Keep original · LF · CRLF
- Include UTF-8 BOM
- Off by default
Capabilities and limits
- Plain-text files are limited to 10 MiB each and 30 MiB total; confirm that legacy-encoding previews are readable.
Merge text filesPut several text files together in the order you choose.
Join local text files in the displayed order. Choose a shared or per-file encoding, review the preview, and enter the exact separator to place between files.
Steps
- Choose the local .txt files in the order you want to process them. The working copy stays in this browser.
- Review Input encoding, Text between files before selecting “Merge text files”.
- Review the TXT result, then use the available copy or download controls.
Available options
- Input encoding
- UTF-8 (strict check) · GB18030 / GBK · Big5
- I checked that the decoded text is readable.
- Off by default
- Text between files
Capabilities and limits
- Plain-text files are limited to 10 MiB each and 30 MiB total; confirm that legacy-encoding previews are readable.
Word counterWords, characters, and paragraphs at a glance.
Count pasted text by Unicode code points, UTF-16 units, grapheme clusters, Latin-letter English words, Han characters, lines, paragraphs and UTF-8 bytes. These measures differ for emoji and mixed-language text; no publisher-specific word count is implied.
Steps
- Paste content into “Text”, or load the example to try it.
- Check your input, then select “Count text”.
- Read the labeled counts and copy a value or the full JSON summary. Compare the required counting rule with your destination's own limit.
Capabilities and limits
- Pasted text is limited to 1 MiB; review the result before using it.
Case converterChange letter case to suit your text.
Change pasted text with Unicode upper/lower mappings or English-oriented title, sentence and surname-first modes. The surname mode makes the first Latin word of each line uppercase and title-cases the remaining English words; review names and acronyms manually.
Steps
- Paste content into “Text”, or load the example to try it.
- Choose lowercase, uppercase, English title or sentence case, or SURNAME Given Name for one name per line; then select “Convert case”.
- Compare the output with the source, especially names and acronyms, then copy or download result.txt.
Available options
- Case style
- lowercase · UPPERCASE · Title Case (English) · Sentence case (English) · SURNAME Given Name (English)
Capabilities and limits
- Pasted text is limited to 1 MiB; review the result before using it.
Remove duplicate linesLess repetition. Same order.
Remove repeated whole lines without changing their order. The first copy stays exactly as written; optional case and surrounding-space rules only affect comparison.
Steps
- Paste one entry per line into “Text”, or load the example.
- Choose whether comparisons ignore letter case or surrounding spaces, then remove duplicates.
- Check the remaining line order and copy or download the complete result.
Available options
- Ignore case when comparing
- Off by default
- Ignore surrounding whitespace when comparing
- Off by default
Capabilities and limits
- Paste up to 1 MiB of text. This compares whole lines, not a selected column or part of a URL; review the result before using it.
- CR, LF and CRLF line breaks use the first detected style in the result. Malformed Unicode is rejected.
Sort text linesSort your lines as text or numbers.
Sort pasted lines as text or exact decimal values, in ascending or descending order. Equal values keep their original order; blank lines go last.
Steps
- Paste one item per line into “Text”, or load the numeric example.
- Choose Unicode lexical order or exact decimal values, then choose ascending or descending.
- Sort, inspect the line order, then copy or download the complete result.
Available options
- Sort as
- Unicode lexical order · Exact decimal values
- Direction
- Ascending · Descending
Capabilities and limits
- Paste up to 1 MiB. Numeric mode requires every nonblank line to be one valid decimal number, with no labels or separators.
- Numeric exponents must be between −1,000,000 and 1,000,000. Text order compares UTF-16 units without locale collation. Malformed Unicode is rejected.
Clean whitespaceTidy extra spaces and empty lines without losing useful breaks.
Clean pasted text while keeping useful line and paragraph breaks. Remove trailing spaces, optionally reduce spaces and tabs inside lines, and choose how many blank lines to retain.
Steps
- Paste text into “Text”, or load the example to try it.
- Choose trailing-space removal, space/tab collapsing, and the maximum consecutive blank lines (0 removes all).
- Review paragraph spacing, then copy or download the complete result.
Available options
- Remove trailing spaces
- On by default
- Collapse spaces and tabs
- Off by default
- Maximum consecutive empty lines
- 1
Capabilities and limits
- Paste up to 1 MiB. This tool does not join wrapped lines into paragraphs; review the spacing before using the result.
- Trailing cleanup removes tabs and Unicode space separators. Space/tab collapsing leaves other characters intact; malformed Unicode is rejected. Mixed CR, LF and CRLF breaks use the first detected style.
Find & replaceFind the text you need and replace it with care.
Replace literal matches by default. Optional JavaScript regular expressions run in a disposable worker with a hard time limit; regex replacement tokens follow JavaScript rules.
Steps
- Paste content into “Text”, or load the example to try it.
- Review Find, Replace with before selecting “Find and replace”.
- Review the Text result, then use the available copy or download controls.
Available options
- Find
- Enter as needed
- Replace with
- Enter as needed
- Replace all matches
- On by default
- Match case
- On by default
- Use a regular expression
- Off by default
- Multiline anchors
- Off by default
Capabilities and limits
- Pasted text is limited to 1 MiB; review the result before using it.
Slug generatorTurn a title into a tidy URL slug.
Normalize text with NFKC, lowercase it, and join letters and numbers with hyphens. Non-Latin letters remain by default. ASCII-only mode strips decomposable accents before removing other non-ASCII characters; it does not transliterate Chinese.
Steps
- Paste content into “Text”, or load the example to try it.
- Check your input, then select “Generate slug”.
- Review the URL slug result, then use the available copy or download controls.
Available options
- ASCII-only slug (é becomes e; Chinese characters are removed)
- Off by default
Capabilities and limits
- Pasted text is limited to 1 MiB; review the result before using it.
Split text fileSplit a long text file by line count or size.
Split text into line-count or UTF-8-byte-bounded parts and download a real ZIP. Byte splitting never cuts a Unicode code point; joining the parts restores the decoded source.
Steps
- Paste content into “TXT” or choose a local file.
- Review Input encoding, Split by, Lines / maximum bytes per part before selecting “Split text”.
- Download the ZIP, inspect the numbered TXT parts, and join their contents in order if you need to reconstruct the decoded source.
Available options
- Input encoding
- UTF-8 (strict check) · GB18030 / GBK · Big5
- I checked that the decoded text is readable.
- Off by default
- Split by
- Line count · UTF-8 byte budget
- Lines / maximum bytes per part
- 1000
Switching modes resets this value to 1,000 lines or 1,048,576 UTF-8 bytes. Byte mode may split a line, but never a Unicode code point.
Capabilities and limits
- A local TXT file is limited to 10 MiB; pasted text can be up to 30 MiB. ZIP archives contain at most 1,000 parts. Confirm that legacy-encoding previews are readable.
Email extractorExtract and deduplicate common email addresses from pasted text or a UTF-8 text file.
Find email addresses in a copied contact list, message or text/HTML source file. Review the distinct addresses, then copy or download the list.
Steps
- Choose Pasted text or Text file. Paste a list, message or HTML snippet, or select one UTF-8 file.
- Choose whether addresses that differ only in letter case should be merged; the first spelling is kept.
- Extract, review the count and preview, then copy or download the full one-address-per-line TXT list.
Capabilities and limits
- Up to 1 MiB pasted text or one UTF-8 text file up to 10 MiB; up to 50,000 lines and 50,000 address matches. HTML and CSV are scanned as text, not parsed into fields.
- Recognizes common ASCII addresses with dot-atom local parts and domain names. Quoted local parts and internationalized addresses are not covered; no mailbox, consent or deliverability check is performed.
URL extractorPull absolute HTTP and HTTPS links from text or a UTF-8 file into a list you can copy or save.
Extract written HTTP and HTTPS links from notes, logs or source text. Review the list, then copy it or download a line-by-line TXT file. The tool does not visit the links.
Steps
- Choose pasted text or one UTF-8 text file, then provide content containing links.
- Choose HTTP, HTTPS or both, then optionally list allowed top-level domains and host domains. Decide whether to keep repeated links.
- Select Extract links, review the first matches, then copy or download the complete list.
Available options
- Protocol
- HTTP and HTTPS · HTTP only · HTTPS only
- Allowed TLDs (optional)
- Enter as needed
Comma-separated final labels, for example com,org. Blank allows all.
- Allowed domains (optional)
- Enter as needed
Comma-separated hosts, for example example.com; subdomains count too. Blank allows all.
Capabilities and limits
- Paste up to 1 MiB or choose one UTF-8 text file up to 10 MiB. At most 50,000 lines and 50,000 matched links are processed locally.
- Only written absolute HTTP/HTTPS links are scanned. Relative HTML hrefs and links hidden behind rich text are not resolved; no link is opened or checked for reachability.
Line editing workspaceWrap text, number or strip lines, add affixes, create lists and extract ranges in one workspace.
Edit text line by line in one workspace. Reflow paragraphs, join soft breaks, add or remove line numbers, add prefixes or suffixes, make Markdown lists, or extract a line range. Choose one task and review the output before copying or downloading.
Steps
- Choose a task. For range extraction, select line numbers or literal start/end text from a static file snapshot.
- Paste text or select one UTF-8 text file.
- Run the task, inspect the result and line counts, then copy or download it.
Available options
- Task
- Wrap / unwrap · Add numbers · Remove numbers · Add affixes · Create list · Extract range
- Line width
- 80
- Mode
- wrap · unwrap
- Range selection
- Line numbers · Text markers
- Start at
- 1
- Separator
- .
- Prefix
- item-
- Suffix
- Enter as needed
- List style
- bullet · numbered · checklist
- Last line (optional)
- Enter as needed
- Start text
- Enter as needed
First line containing this exact text is included.
- End text (optional)
- Enter as needed
First later line containing this exact text is included; blank means through the end.
Capabilities and limits
- Wrap joins lines within each paragraph before reflowing; blank lines separate paragraphs and long words stay intact.
- Numbering, affixes and lists preserve a final newline without adding an extra item. Extracted ranges omit a trailing newline; blank list lines stay blank.
- Paste up to 1 MiB or choose one UTF-8 file up to 10 MiB, with at most 50,000 physical lines. Full output is limited to 20 MiB; all option values share a 2 MiB budget. Malformed Unicode in text or active literal options is rejected.
- Wrapping counts Unicode code points, rather than display cells or combined-character clusters. Files use strict UTF-8 and discard an initial BOM; output line breaks use LF.
Word and character frequencyRank words or Unicode code points and locate a chosen word or phrase in the source text.
Paste a draft or open a UTF-8 text file to rank its words or Unicode code points. In word mode, enter a known phrase to see its count and first 20 source locations. Rankings count exact word forms; the tool does not judge writing quality.
Steps
- Choose Words or Code points. To locate a known phrase, enter it in word mode.
- Paste text or choose one UTF-8 file, then count frequencies.
- Inspect the ranking and phrase locations, copy matching ranking rows, or download the complete JSON/CSV ranking.
Available options
- Task
- Word groups · Code points
- Hide common English function words
- Off by default
- Include spaces and line breaks
- On by default
- Find a word or phrase in the source (optional)
- Enter as needed
Up to 120 characters. Exact case-insensitive match; the first 20 source locations are shown.
Capabilities and limits
- Word rankings ignore case after Unicode NFC normalization. Phrase lookup is an exact case-insensitive match in the source with flexible whitespace; related forms are not grouped.
- Character mode counts Unicode code points, including whitespace by default. One visible symbol can contain multiple code points.
- Pasted text is limited to 1 MiB; one UTF-8 file to 10 MiB; at most 50,000 physical lines and 50,000 distinct values. All processing stays in this browser.
Compare two listsFind shared entries, items only in A or B, and a combined list.
Compare two plain-text lists, one item per line. See A-only, shared and B-only groups with each preview item's first source line, so you can find it in the original list.
Steps
- Paste list A or choose a UTF-8 file, then paste list B. Put one item on each line.
- Choose the output group and decide whether case and surrounding spaces should match.
- Inspect all three groups and first source lines, then copy a complete group or download the selected TXT.
Available options
- Download / preview
- Shared · in A and B · A only · absent from B · B only · absent from A · Only in either · A-only + B-only · Combined · every distinct item
- Ignore letter case
- On by default
- Ignore spaces around items
- On by default
Capabilities and limits
- Paste up to 1 MiB per list or choose one UTF-8 file for A (up to 10 MiB); each list is limited to 50,000 lines.
- Blank lines and duplicates are skipped. This compares whole lines, not CSV columns, Excel workbooks or fuzzy names, and does not highlight the original sheet.
- All processing happens in this browser; no account or upload.
Unicode text normalizerNormalize Unicode text to NFC, NFD, NFKC or NFKD and optionally compare two texts after normalization.
Text can look identical yet use different Unicode code points. Paste text or choose one UTF-8 file, optionally add a second text, and check whether they become identical in the selected form before copying or downloading the result.
Steps
- Paste text or choose one UTF-8 file. For two filenames or strings, paste the other one into the optional second input.
- Choose NFC or NFD for canonical equivalence; use NFKC/NFKD only when compatibility folding is intentional.
- Inspect the first code point change and, if supplied, the before/after comparison. Copy or download the normalized first text as UTF-8 TXT.
Available options
- Normalization form
- NFC · compose (general text) · NFD · decompose · NFKC · compatibility + compose · NFKD · compatibility + decompose
Capabilities and limits
- Pasted text and the optional second text are each limited to 1 MiB; one UTF-8 file is limited to 10 MiB and 50,000 physical lines. Input stays in this browser.
- NFKC/NFKD can replace compatibility characters and lose distinctions. Keep the original if spelling, symbols or formatting matter.
- Normalization does not strip accents, transliterate scripts, remove invisible characters or check visually confusable text.
Clean selected hidden charactersFind zero-width spaces and other selected hidden code points, then remove or replace only the types you choose.
Diagnose hidden characters in pasted text or a UTF-8 file. The default removes only zero-width spaces; other characters stay unless you select them. Review counts and source positions before using the result.
Steps
- Paste text or choose one UTF-8 text file.
- Choose exactly which hidden characters to remove. Optionally turn non-breaking spaces into ordinary spaces.
- Run the tool, inspect the per-character counts and first source positions, then copy or download the cleaned TXT.
Available options
- Remove zero-width spaces · U+200B
- On by default
- Remove word joiners · U+2060
- Off by default
- Remove FEFF characters · U+FEFF
- Off by default
- Remove ZWNJ / ZWJ · U+200C / U+200D
- Off by default
- Replace non-breaking spaces with spaces · U+00A0
- Off by default
- Remove left-to-right marks · U+200E
- Off by default
Capabilities and limits
- Only the seven listed code points are checked; this is not a general Unicode sanitizer.
- U+200E can affect text direction; ZWJ/ZWNJ can shape scripts and emoji. Check the original before removing them.
- Pasted text is limited to 1 MiB; one UTF-8 text file is limited to 10 MiB and 50,000 physical lines. An initial file BOM may be consumed during UTF-8 decoding.
Unicode inspectorInspect each Unicode code point, its formal name or control alias, UTF-8 bytes, and source position.
Paste text or choose one UTF-8 text file to identify the exact code points behind a visible string. Search every result by code, name or bytes; export the complete inspection as JSON or CSV.
Steps
- Paste the text or select one UTF-8 text file.
- Run the inspector, then search the full result by code point, Unicode name or UTF-8 bytes. Check each source line and column.
- Copy matching rows as JSON or copy the full JSON, or download the complete JSON or CSV report.
Capabilities and limits
- Up to 10,000 Unicode code points per run. Pasted text is limited to 1 MiB; one UTF-8 .txt/.md file is limited to 10 MiB and 50,000 physical lines.
- Names come from Unicode 18.0.0. Controls use Unicode control aliases; private-use and unassigned code points have no formal name. An initial file BOM may be consumed during UTF-8 decoding.
- Rows describe individual code points, not complete grapheme clusters or visual-confusable equivalence. Input is processed in this browser.