Check an original word list with an explicit apostrophe policy
Prepare complete TXT or JSON words, choose the active policy, inspect every row and keep the full export package together.
Identify the complete original
Each TXT line is a distinct word; a trailing delimiter leaves an empty word. JSON must contain one top-level string array. Duplicates, empty words, whitespace and original escaped lexemes remain. Numbers or booleans are never coerced into words. An initial UTF-8 BOM stays in the original and is excluded from the first word.
["hello","aren’t","aren't"," hello ","zzz21stzzz","-0","","hello"]- Select one complete UTF-8 file or pasted word list and separately choose TXT lines or a JSON string array for words and the optional whitelist.
- Select the active apostrophe policy. The whitelist matches original characters and case exactly; an empty string overrides only an exact empty word.
Select a reproducible policy
Native dictionary policy can convert curly apostrophes, so it is not literal equality. Original-apostrophe policy still trims and applies case rules. Settings record all three definitions and the selected policy; correct/incorrect alone does not explain what happened.
| Policy | Actual lookup |
|---|---|
| native-dictionary-policy | Native trim, ICONV, case, affix and compound rules |
| tool-curly-policy | Only U+2019 → U+0027 before native rules |
| original-apostrophe-policy | Disable this job’s ICONV; native trim, case, affix and compound rules remain |
Audit each row and retain the whole package
The complete table has 26 columns. Ordinals, source positions and rawToken distinguish duplicates. Transformations and actual candidate attempts come from the same correct call. A whitelisted row has no dictionary call, so null and false mean different things. Complete output and independent readback of this example have been checked; UI and native Worker acceptance remain pending.
| Files | Purpose |
|---|---|
| spelling-results.csv / spelling-report.json | All rows, traces and budget snapshot |
| words-original.* / whitelist-original.* | Byte-exact originals; whitelist original exists only when enabled |
| dictionary-en-US.aff / .dic / whitelist-active.json | Active dictionary and every whitelist group |
| spelling-settings.json / *LICENSE.txt | Fixed settings, source hashes and three complete licenses |
Limits, attribution and recovery
Words plus whitelist share 4 MiB UTF-8; words have at most 100,000 rows and the whitelist separately has at most 100,000 entries; each word has at most 128 Unicode codepoints.
The fixed AFF and DIC share 8 MiB; at most 100,000 word checks; loading, expansion, whitelist, lookup and serialization share 100,000,000 observable charged work units.
All 10/11 files plus independent complete copy text share 64 MiB. Held sources, dictionary, rows, serialization, clones, Blobs and UI representations share a conservative 768 MiB preallocation guard.
The absolute 60-second deadline starts at input capture and covers reading, module loading, dictionary initialization, checking, encoding, complete validation, cleanup and first publication; no pre-expanded dictionary bypass.
Regex VM internal steps, native container memory estimates, joint capacity, complete Hunspell equivalence, native Worker and cross-engine acceptance remain pending. The original private word list and environment were not published. This is not grammar certification.
Dictionary includes UKACD, SCOWL, WordNet, VarCon and Ispell notices. Read all licenses in the supplied documentation.
Check again with the same sources and settings
- Changing input or policy immediately invalidates old results; do not mix downloads from different runs.
- The original private word list was not published; this tool cannot infer correctness for that entire corpus.
References
- nspell 2.1.5 source
Pinned mature lookup branches; this tool marks budget observation and whole-word anchor changes.
- dictionary-en source and full licenses
Pinned dictionary-en 4.0.0; full upstream notices are exported. The linked current branch does not replace the pinned version.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Original word list spelling checkCheck every original English word under an explicit apostrophe policy and retain complete lookup and whitelist evidence.
Choose one of three apostrophe policies for the fixed en_US dictionary. Repeated words, whitespace, empty words and original JSON lexemes keep distinct ordinals and source positions. An exact-original whitelist can override a word without changing the dictionary.
Steps
- Select one complete UTF-8 file or pasted word list and separately choose TXT lines or a JSON string array for words and the optional whitelist.
- Select the active apostrophe policy. The whitelist matches original characters and case exactly; an empty string overrides only an exact empty word.
- Read all 26 columns and actual transformation/lookup traces across every page. Use complete-file readers for originals, settings, dictionary and all licenses.
- Copy the complete report and inventory and save all 10 files, or 11 files including the whitelist original when enabled.
Available options
- Word source
- Local file · Paste complete text
- Word-list format
- TXT: one word per line · JSON: string array
- Apostrophe policy
- Dictionary policy: includes native curly-apostrophe conversion · Convert U+2019 to ASCII apostrophe in the lookup copy · Keep original apostrophes; disable dictionary input conversion
- Exact original-word whitelist
- No whitelist · Local file · Paste complete text
- Word-list format
- TXT: one word per line · JSON: string array
Capabilities and limits
- Words plus whitelist share 4 MiB UTF-8; words have at most 100,000 rows and the whitelist separately has at most 100,000 entries; each word has at most 128 Unicode codepoints.
- The fixed AFF and DIC share 8 MiB; at most 100,000 word checks; loading, expansion, whitelist, lookup and serialization share 100,000,000 observable charged work units.
- All 10/11 files plus independent complete copy text share 64 MiB. Held sources, dictionary, rows, serialization, clones, Blobs and UI representations share a conservative 768 MiB preallocation guard.
- The absolute 60-second deadline starts at input capture and covers reading, module loading, dictionary initialization, checking, encoding, complete validation, cleanup and first publication; no pre-expanded dictionary bypass.
- TXT retains CRLF, CR, LF and trailing empty rows; JSON accepts only a complete string array. Invalid UTF-8 and isolated surrogates are rejected. String "-0" stays text; a numeric -0 word is rejected.
- Regex VM internal steps, native container memory estimates, joint capacity, complete Hunspell equivalence, native Worker and cross-engine acceptance remain pending. The original private word list and environment were not published. This is not grammar certification.