Multilocus sequence-matrix builder
Join already aligned DNA FASTA loci by exact taxon ID and export a complete FASTA/NEXUS supermatrix, locus partitions, presence map and original sources.
- 1Add input
- 2Adjust settings
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Select ordered aligned loci. Union explicitly fills absent taxa; intersection keeps only shared taxa. This tool concatenates existing alignments and creates charsets.
How to use this tool
- Select the loci in the desired order; each file must already contain an alignment.
- Choose union/intersection and the missing-locus symbol, then check taxa, columns, filled cells and presence rows.
- Download complete concatenated.fasta, concatenated.nex, presence.csv, report.json and all original files.
Supported inputs and limits
Up to20 files with10MiB total,5,000 output taxa,500,000 total columns,10,000,000 matrix cells,100,000 source records and300,000 physical lines. Each line≤65,536 bytes and header≤4,096 bytes. Complete derived files and originals≤48MiB.
Limits apply together:5,000 taxa ×500,000 columns exceeds the10million-cell guard. The source-record maximum follows20 loci ×5,000 unique source taxa; it is not an additional independent100,001-record supported boundary.
UTF-8 with optional BOM, LF/CRLF, aligned DNA IUPAC symbols plus ? and -. Sequences must be equally long within each locus; IDs are the case-sensitive first header token and unique within a locus. Full headers, source spans, case and original bytes stay intact.
RNA/protein, comments and embedded sequence spaces are Unsupported. No aligner, fuzzy species-name matching or phylogenetic inference. Output taxon order follows first occurrence in source-file/record order; file order defines locus columns.
NEXUS exports DNA matrix and fixed locus_ordinal charsets with complete original file labels in JSON. Every taxon/locus presence or filled state appears in presence.csv/report.json. Formula-like CSV taxon labels are protected with a leading apostrophe; unchanged labels remain in JSON.
Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
Worked example
Example input
locus A: 44 bytes; locus B: 26 bytes
Example options
{"params":{"missing":"?","taxaMode":"union"}}Example output
{"schemaVersion":1,"profile":"aligned-DNA-FASTA-first-token-identifiers","taxaMode":"union","missing":"?","order":"first occurrence across selected file order and source record order","taxa":["a","b","c"],"columns":7,"matrixCells":21,"missingCells":7,"inputFiles":2,"inputBytes":70,"sourceRecords":4,"physicalLines":8,"partitions":[{"ordinal":1,"name":"locus A","nexusName":"locus_1","start":1,"end":4,"length":4},{"ordinal":2,"name":"locus B","nexusName":"locus_2","start":5,"end":7,"length":3}],"loci":[{"ordinal":1,"name":"locus A","length":4,"records":[{"id":"a","title":"a preserved title","headerLine":1,"headerSpan":[0,18],"sequenceSpans":[[19,23]],"residueCount":4},{"id":"b","title":"b description","headerLine":3,"headerSpan":[24,38],"sequenceSpans":[[39,43]],"residueCount":4}]},{"ordinal":2,"name":"locus B","length":3,"records":[{"id":"b","title":"b desc2","headerLine":1,"headerSpan":[0,8],"sequenceSpans":[[9,12]],"residueCount":3},{"id":"c","title":"c title","headerLine":3,"headerSpan":[13,21],"sequenceSpans":[[22,25]],"residueCount":3}]}],"presence":[[true,false],[true,true],[false,true]],"taxonDisposition":[{"id":"a","inOutput":true},{"id":"b","inOutput":true},{"id":"c","inOutput":true}],"normalization":"Residue case preserved; input line endings/header descriptions retained in original files. Missing loci use the explicit chosen symbol."}When something does not work
Check the selected profile and the original source, then correct the input or reduce complete input/output size. Cancellation and errors publish no partial files; rerun the same supported source.
Frequently asked questions
Does it align sequences?
No. It joins existing equal-length locus alignments by exact case-sensitive first-token IDs.
Are absent loci silently removed?
No. Union fills them with the selected symbol and records presence; intersection deliberately retains only taxa found in all loci.