Neatbo.

Multilocus sequence-matrix builder

Join already aligned DNA FASTA loci by exact taxon ID and export a complete FASTA/NEXUS supermatrix, locus partitions, presence map and original sources.

Browser-local processingInputAligned DNA FASTA filesOutputFASTA / NEXUS / CSV / JSON / originalsUp to 10 MiB per file · File limit: 20
  1. 1Add input
  2. 2Adjust settings
  3. 3Get your result

Tool input and files are processed in this browser without being uploaded.

Your input

Inputs are kept temporarily in this tab when switching tools. Refreshing or closing clears them; large results may need to be regenerated.

⌘ / Ctrl + Enter to run

or drag and drop them here

Files stay on this device. Your originals stay unchanged.

Up to 10 MiB per file · File limit: 20

    Options

    Complete the required options first. You can keep the defaults for the rest.

    Preparing the tool…

    Before you start

    Select ordered aligned loci. Union explicitly fills absent taxa; intersection keeps only shared taxa. This tool concatenates existing alignments and creates charsets.

    How to use this tool

    1. Select the loci in the desired order; each file must already contain an alignment.
    2. Choose union/intersection and the missing-locus symbol, then check taxa, columns, filled cells and presence rows.
    3. Download complete concatenated.fasta, concatenated.nex, presence.csv, report.json and all original files.

    Supported inputs and limits

    Up to20 files with10MiB total,5,000 output taxa,500,000 total columns,10,000,000 matrix cells,100,000 source records and300,000 physical lines. Each line≤65,536 bytes and header≤4,096 bytes. Complete derived files and originals≤48MiB.

    Limits apply together:5,000 taxa ×500,000 columns exceeds the10million-cell guard. The source-record maximum follows20 loci ×5,000 unique source taxa; it is not an additional independent100,001-record supported boundary.

    UTF-8 with optional BOM, LF/CRLF, aligned DNA IUPAC symbols plus ? and -. Sequences must be equally long within each locus; IDs are the case-sensitive first header token and unique within a locus. Full headers, source spans, case and original bytes stay intact.

    RNA/protein, comments and embedded sequence spaces are Unsupported. No aligner, fuzzy species-name matching or phylogenetic inference. Output taxon order follows first occurrence in source-file/record order; file order defines locus columns.

    NEXUS exports DNA matrix and fixed locus_ordinal charsets with complete original file labels in JSON. Every taxon/locus presence or filled state appears in presence.csv/report.json. Formula-like CSV taxon labels are protected with a leading apostrophe; unchanged labels remain in JSON.

    Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.

    Worked example

    Example input

    locus A: 44 bytes; locus B: 26 bytes
    Example options
    {"params":{"missing":"?","taxaMode":"union"}}

    Example output

    {"schemaVersion":1,"profile":"aligned-DNA-FASTA-first-token-identifiers","taxaMode":"union","missing":"?","order":"first occurrence across selected file order and source record order","taxa":["a","b","c"],"columns":7,"matrixCells":21,"missingCells":7,"inputFiles":2,"inputBytes":70,"sourceRecords":4,"physicalLines":8,"partitions":[{"ordinal":1,"name":"locus A","nexusName":"locus_1","start":1,"end":4,"length":4},{"ordinal":2,"name":"locus B","nexusName":"locus_2","start":5,"end":7,"length":3}],"loci":[{"ordinal":1,"name":"locus A","length":4,"records":[{"id":"a","title":"a preserved title","headerLine":1,"headerSpan":[0,18],"sequenceSpans":[[19,23]],"residueCount":4},{"id":"b","title":"b description","headerLine":3,"headerSpan":[24,38],"sequenceSpans":[[39,43]],"residueCount":4}]},{"ordinal":2,"name":"locus B","length":3,"records":[{"id":"b","title":"b desc2","headerLine":1,"headerSpan":[0,8],"sequenceSpans":[[9,12]],"residueCount":3},{"id":"c","title":"c title","headerLine":3,"headerSpan":[13,21],"sequenceSpans":[[22,25]],"residueCount":3}]}],"presence":[[true,false],[true,true],[false,true]],"taxonDisposition":[{"id":"a","inOutput":true},{"id":"b","inOutput":true},{"id":"c","inOutput":true}],"normalization":"Residue case preserved; input line endings/header descriptions retained in original files. Missing loci use the explicit chosen symbol."}

    When something does not work

    Check the selected profile and the original source, then correct the input or reduce complete input/output size. Cancellation and errors publish no partial files; rerun the same supported source.

    Frequently asked questions

    Does it align sequences?

    No. It joins existing equal-length locus alignments by exact case-sensitive first-token IDs.

    Are absent loci silently removed?

    No. Union fills them with the selected symbol and records presence; intersection deliberately retains only taxa found in all loci.

    Documentation & further reading

    Related tools