Aligned FASTA whole-column trimming
Remove alignment columns using an explicit missing-character set and threshold, preserving every row, full header and complete original-to-output column map.
- 1Add input
- 2Adjust settings
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Apply one column decision to all equal-width sequences. Inspect missing counts and retained coordinates before downloading the projected alignment; ambiguous N/X characters count only when explicitly selected.
How to use this tool
- Select or paste one equal-width aligned FASTA. Choose the explicit missing-character set and column rule.
- Run trimming and inspect original, retained and removed column counts. Check the complete column map when positions matter.
- Download projected FASTA and full JSON/CSV. Keep the original source and coordinates with the handoff.
Supported inputs and limits
One equal-width UTF-8 aligned FASTA up to 10 MiB, with optional BOM and LF/CRLF. Up to 10,000 rows, 100,000 source columns and 8,000,000 total residue cells; full headers up to 65,536 UTF-16 units and at most 200,000 nonempty physical sequence fragments. Complete JSON, CSV and FASTA together must fit 40 MiB.
Residues are ASCII A–Z/a–z, *, ?, . or -. Full UTF-8 headers, duplicate IDs and source spans remain distinct. Lone CR, unsupported residues, unequal widths and control characters in headers reject; blank source lines are retained in the original text.
Choose missing characters explicitly, then any missing value or an integer 0–100% rule. Percentage removes a column when missing count ×100 is greater than or equal to row count ×threshold. At90%,9 of10 removes and8 of10 remains; at0%, every column is removed.
N/X ambiguity is not a gap unless you select it. This tool does not align sequences, infer relationships or choose biologically optimal trimming. The cited220MB dataset is outside the10MiB profile.
Full JSON retains original source/hash, every full header, physical source spans, all missing counts, the complete zero-based column map and every projected sequence. CSV maps every source column; null/new empty coordinate means removed. No dense lists of millions of cause indices are silently truncated.
Zero-column output is valid and explicit: every original header has an empty sequence. Output FASTA uses LF; original LF/CRLF/BOM remain reconstructable from report source text. Preview is200 rows/2,000 UTF-16 units per cell; large copies are labelled previews while complete downloads remain atomic.
Worked example
Example input
>row1 A-- >row2 A-- >row3 A-- >row4 A-- >row5 A-- >row6 A-- >row7 A-- >row8 A-- >row9 A-C >row10 ACC
Example options
{"secondary":"","params":{"missingCharacters":"-","rule":"percentage","percent":90,"spreadsheetSafe":true}}Example output
{"source":{"name":"pasted.txt","bytes":101,"sha256":"989e4f0278a17af599cc19fb61849719d8910618570894ee9301336b3deab116","encoding":"UTF-8","bomRetained":false},"sourceText":">row1\nA--\n>row2\nA--\n>row3\nA--\n>row4\nA--\n>row5\nA--\n>row6\nA--\n>row7\nA--\n>row8\nA--\n>row9\nA-C\n>row10\nACC\n","summary":{"rows":10,"inputColumns":3,"outputColumns":2,"removedColumns":1,"inputCells":30,"maxHeaderUTF16":5,"physicalSequenceFragments":10,"status":"aligned-columns-projected"},"missingCharacters":["-"],"threshold":90,"comparison":"greater-than-or-equal-integer-cross-product","coordinates":"zero-based","causes":"complete per-column missing count; raw source determines each contributing row","missingCounts":[0,9,8],"columnMap":[0,null,1],"records":[{"header":"row1","sequence":"A-","sourceHeader":{"utf8Start":1,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":6,"utf8Bytes":3,"line":2}]},{"header":"row2","sequence":"A-","sourceHeader":{"utf8Start":11,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":16,"utf8Bytes":3,"line":4}]},{"header":"row3","sequence":"A-","sourceHeader":{"utf8Start":21,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":26,"utf8Bytes":3,"line":6}]},{"header":"row4","sequence":"A-","sourceHeader":{"utf8Start":31,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":36,"utf8Bytes":3,"line":8}]},{"header":"row5","sequence":"A-","sourceHeader":{"utf8Start":41,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":46,"utf8Bytes":3,"line":10}]},{"header":"row6","sequence":"A-","sourceHeader":{"utf8Start":51,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":56,"utf8Bytes":3,"line":12}]},{"header":"row7","sequence":"A-","sourceHeader":{"utf8Start":61,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":66,"utf8Bytes":3,"line":14}]},{"header":"row8","sequence":"A-","sourceHeader":{"utf8Start":71,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":76,"utf8Bytes":3,"line":16}]},{"header":"row9","sequence":"AC","sourceHeader":{"utf8Start":81,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":86,"utf8Bytes":3,"line":18}]},{"header":"row10","sequence":"AC","sourceHeader":{"utf8Start":91,"utf8Bytes":5},"sourceSequenceSpans":[{"utf8Start":97,"utf8Bytes":3,"line":20}]}],"scope":"One declared equal-width alignment; ASCII residues [A-Za-z*?.-], full UTF8 headers and duplicate IDs preserved. N/X count only if explicitly selected. No alignment, inference or biological optimality. 220MB source use case is outside the10MiB input profile."}When something does not work
Correct unequal widths or unsupported residues and confirm your missing-character set and integer threshold. Reduce source, cells, fragments or complete-output size after a limit rejection; preserve the original alignment. Cancellation publishes no partial FASTA or map.
Frequently asked questions
Are gaps removed independently from each sequence?
No. One shared column mask is applied to every row, preserving alignment. The starter removes the90%-missing column and keeps the80%-missing column.
Do N and X count as missing by default?
No. The default set is -. Add N/X only if your receiving task explicitly defines them as missing; that decision changes the shared column mask.
Can every output column disappear?
Yes. The result clearly reports zero output columns and writes each full header with an empty sequence. At0%, the >= rule removes every column. Confirm the threshold before downstream use.
Documentation & further reading
Related tools
JSON formatting workspace
Format or minify strict JSON, sort object keys, and encode or decode strings while preserving raw number tokens.
Regex tester
Try a pattern and see what it matches in your text.
Compare text
See what changed, side by side.
HTML formatter
Format HTML indentation so its structure is easier to read.