Neatbo.

Aligned FASTA whole-column trimming

Remove alignment columns using an explicit missing-character set and threshold, preserving every row, full header and complete original-to-output column map.

Browser-local processingInputAligned UTF-8 FASTAOutputFASTA / map JSON / CSVUp to 10 MiB per file · File limit: 1
  1. 1Add input
  2. 2Adjust settings
  3. 3Get your result

Tool input and files are processed in this browser without being uploaded.

Your input

Inputs are kept temporarily in this tab when switching tools. Refreshing or closing clears them; large results may need to be regenerated.

⌘ / Ctrl + Enter to run

or drag and drop it here

Files stay on this device. Your originals stay unchanged.

.fa · .fasta · .fas · .aln · .txt

Up to 10 MiB per file · File limit: 1

    0 characters · 0 bytes
    Options

    Complete the required options first. You can keep the defaults for the rest.

    Preparing the tool…

    Before you start

    Apply one column decision to all equal-width sequences. Inspect missing counts and retained coordinates before downloading the projected alignment; ambiguous N/X characters count only when explicitly selected.

    How to use this tool

    1. Select or paste one equal-width aligned FASTA. Choose the explicit missing-character set and column rule.
    2. Run trimming and inspect original, retained and removed column counts. Check the complete column map when positions matter.
    3. Download projected FASTA and full JSON/CSV. Keep the original source and coordinates with the handoff.

    Supported inputs and limits

    One equal-width UTF-8 aligned FASTA up to 10 MiB, with optional BOM and LF/CRLF. Up to 10,000 rows, 100,000 source columns and 8,000,000 total residue cells; full headers up to 65,536 UTF-16 units and at most 200,000 nonempty physical sequence fragments. Complete JSON, CSV and FASTA together must fit 40 MiB.

    Residues are ASCII A–Z/a–z, *, ?, . or -. Full UTF-8 headers, duplicate IDs and source spans remain distinct. Lone CR, unsupported residues, unequal widths and control characters in headers reject; blank source lines are retained in the original text.

    Choose missing characters explicitly, then any missing value or an integer 0–100% rule. Percentage removes a column when missing count ×100 is greater than or equal to row count ×threshold. At90%,9 of10 removes and8 of10 remains; at0%, every column is removed.

    N/X ambiguity is not a gap unless you select it. This tool does not align sequences, infer relationships or choose biologically optimal trimming. The cited220MB dataset is outside the10MiB profile.

    Full JSON retains original source/hash, every full header, physical source spans, all missing counts, the complete zero-based column map and every projected sequence. CSV maps every source column; null/new empty coordinate means removed. No dense lists of millions of cause indices are silently truncated.

    Zero-column output is valid and explicit: every original header has an empty sequence. Output FASTA uses LF; original LF/CRLF/BOM remain reconstructable from report source text. Preview is200 rows/2,000 UTF-16 units per cell; large copies are labelled previews while complete downloads remain atomic.

    Worked example

    Example input

    >row1
    A--
    >row2
    A--
    >row3
    A--
    >row4
    A--
    >row5
    A--
    >row6
    A--
    >row7
    A--
    >row8
    A--
    >row9
    A-C
    >row10
    ACC
    
    Example options
    {"secondary":"","params":{"missingCharacters":"-","rule":"percentage","percent":90,"spreadsheetSafe":true}}

    Example output

    {"source":{"name":"pasted.txt","bytes":101,"sha256":"989e4f0278a17af599cc19fb61849719d8910618570894ee9301336b3deab116","encoding":"UTF-8","bomRetained":false},"sourceText":">row1\nA--\n>row2\nA--\n>row3\nA--\n>row4\nA--\n>row5\nA--\n>row6\nA--\n>row7\nA--\n>row8\nA--\n>row9\nA-C\n>row10\nACC\n","summary":{"rows":10,"inputColumns":3,"outputColumns":2,"removedColumns":1,"inputCells":30,"maxHeaderUTF16":5,"physicalSequenceFragments":10,"status":"aligned-columns-projected"},"missingCharacters":["-"],"threshold":90,"comparison":"greater-than-or-equal-integer-cross-product","coordinates":"zero-based","causes":"complete per-column missing count; raw source determines each contributing row","missingCounts":[0,9,8],"columnMap":[0,null,1],"records":[{"header":"row1","sequence":"A-","sourceHeader":{"utf8Start":1,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":6,"utf8Bytes":3,"line":2}]},{"header":"row2","sequence":"A-","sourceHeader":{"utf8Start":11,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":16,"utf8Bytes":3,"line":4}]},{"header":"row3","sequence":"A-","sourceHeader":{"utf8Start":21,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":26,"utf8Bytes":3,"line":6}]},{"header":"row4","sequence":"A-","sourceHeader":{"utf8Start":31,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":36,"utf8Bytes":3,"line":8}]},{"header":"row5","sequence":"A-","sourceHeader":{"utf8Start":41,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":46,"utf8Bytes":3,"line":10}]},{"header":"row6","sequence":"A-","sourceHeader":{"utf8Start":51,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":56,"utf8Bytes":3,"line":12}]},{"header":"row7","sequence":"A-","sourceHeader":{"utf8Start":61,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":66,"utf8Bytes":3,"line":14}]},{"header":"row8","sequence":"A-","sourceHeader":{"utf8Start":71,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":76,"utf8Bytes":3,"line":16}]},{"header":"row9","sequence":"AC","sourceHeader":{"utf8Start":81,"utf8Bytes":4},"sourceSequenceSpans":[{"utf8Start":86,"utf8Bytes":3,"line":18}]},{"header":"row10","sequence":"AC","sourceHeader":{"utf8Start":91,"utf8Bytes":5},"sourceSequenceSpans":[{"utf8Start":97,"utf8Bytes":3,"line":20}]}],"scope":"One declared equal-width alignment; ASCII residues [A-Za-z*?.-], full UTF8 headers and duplicate IDs preserved. N/X count only if explicitly selected. No alignment, inference or biological optimality. 220MB source use case is outside the10MiB input profile."}

    When something does not work

    Correct unequal widths or unsupported residues and confirm your missing-character set and integer threshold. Reduce source, cells, fragments or complete-output size after a limit rejection; preserve the original alignment. Cancellation publishes no partial FASTA or map.

    Frequently asked questions

    Are gaps removed independently from each sequence?

    No. One shared column mask is applied to every row, preserving alignment. The starter removes the90%-missing column and keeps the80%-missing column.

    Do N and X count as missing by default?

    No. The default set is -. Add N/X only if your receiving task explicitly defines them as missing; that decision changes the shared column mask.

    Can every output column disappear?

    Yes. The result clearly reports zero output columns and writes each full header with an empty sequence. At0%, the >= rule removes every column. Confirm the threshold before downstream use.

    Documentation & further reading

    Related tools