Local HMMER, MARC and supermatrix workflows
Export complete search hits and catalogue fields, or join aligned DNA loci, with explicit profiles, source spans and bounded full downloads.
Choose the receiving task
HMMER export reads already produced search hits; MARC export reads catalogue exchange bytes; supermatrix construction joins existing DNA alignments by taxon. They require different structure, even when the final handoff includes a table.
Select actual profiles before processing. A selected file takes precedence over pasted text; MARC requires a raw file and supermatrix requires ordered files. File extensions do not certify content. Processing stays local and does not execute search programs, catalogue scripts or aligners.
| Task | Derived data | Preserved source |
|---|---|---|
| HMMER | report.json + records.csv | original.txt; every line/description span |
| MARC | records.json + fields.csv | original.mrc; every ordered field/subfield/value span |
| Supermatrix | FASTA + NEXUS + presence.csv + report.json | all original loci; full headers and source-record spans |
HMMER: keep accession and program roles
Choose tblout with18 fixed fields or domtblout with22, and explicitly choose hmmscan, hmmsearch or phmmer. The remaining description is free text, not another fixed-width column. Statistical tokens such as1e-400 stay lexical text and are never rounded by the tool.
The native ASCII-space field profile rejects unsupported fixed-field TAB, Unicode-space identifiers and ambiguous hash-prefixed target lines. Description TAB/Unicode, comments, BOM and LF/CRLF remain intact. Domain coordinates are raw1-based inclusive; model and sequence roles depend on the selected program. Header contradictions do not silently change that selection.
MARC: encoding and directory order are separate facts
Leader a declares UTF-8. An empty encoding slot can use an explicit known-UTF-8 override, which stays unverified; it does not convert MARC8 or repair unknown bytes. Framing requires24-byte ASCII leaders,12-byte directories,4500 entry maps, two indicators and two-byte subfield identifiers.
Repeated tags and subfields stay ordered. Directory indices and physical indices are both retained, so physically reordered data does not flatten into a different meaning. Catalogue ID/005/order diagnostics are visible, while complete catalogue semantics and character repertoire are not certified.
Supermatrix: join taxa, then expose missing loci
Each input locus must already be aligned DNA IUPAC FASTA with equal-length sequences and unique case-sensitive first-token IDs. Full raw titles remain available. Choose union with ?/-/N absent-locus filling or intersection of all loci; no name normalization or realignment occurs.
Source-file order defines concatenated columns; first taxon occurrence defines row order. NEXUS includes fixed locus_ordinal charsets, with complete original labels and column ranges in JSON. Presence CSV distinguishes present from filled cells, rather than pretending absent sequence was measured.
Keep complete artifacts when preview ends
A small input can expand into a large complete JSON/CSV/matrix. All budgets apply together and an excessive combined output rejects atomically. Preview is limited to200 rows and2,000 UTF-16 units per long cell; large report copy is a compact preview, not every record.
| Task | Input/structure | Full output |
|---|---|---|
| HMMER | 10MiB;100k hits;300k lines;65,536B/line | 48MiB |
| MARC | 10MiB;10k records;100k fields;200k subfields | 64MiB |
| Supermatrix | 20 files/10MiB;5k taxa;500k columns;10M cells;100k source records | 48MiB |
- Keep original bytes beside source spans; normalized strings cannot recreate framing or BOM.
- MARC CSV marks protective apostrophes explicitly. HMMER/presence CSV protected text can be checked against unchanged JSON values; do not guess whether a source apostrophe was synthetic.
- Supermatrix5,000×500,000 is outside the10million-cell budget; file/taxon/column maxima are not an arbitrary simultaneous promise.
- Use supported originals for cancellation and same-source recovery; errors do not provide partial downloads.
References
- HMMER export task
First-person tblout-to-DataFrame task; query accession must not disappear.
- MARC Unicode export task
First-person Czech MARC-to-CSV encoding problem; original linked file was not obtained.
- Multilocus workflow feedback
Researcher describes missing taxa and charset preparation. Their1500-locus speed statement concerns other software.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
HMMER result-table exportExport every HMMER3 protein-table hit with query and target accessions, original numeric tokens, free-text descriptions and exact source byte spans.
Choose the producing program and table type, then inspect complete hits without rerunning HMMER or changing statistical tokens.
Steps
- Choose tblout/domtblout and the actual producing program; paste or select one complete file.
- Check accessions, descriptions and domain coordinate roles; an empty hit table is distinct from a parse error.
- Download complete report.json, records.csv and original.txt.
Available options
- Table type
- tblout —18 fixed fields · domtblout —22 fixed fields
- Producing program
- hmmscan · hmmsearch · phmmer
Capabilities and limits
- One10MiB UTF-8 protein HMMER3 tblout or domtblout;100,000 hit records;300,000 physical lines;65,536 bytes per physical line. Complete original+JSON+CSV must fit48MiB. Budgets apply together.
- Explicit hmmscan/hmmsearch/phmmer and18/22-field table profiles. Header program/version contradictions, HMMER2/nhmmer, fixed-field TAB, Unicode-space identifiers and ambiguous hash-prefixed target rows are Unsupported. ASCII spaces separate fixed fields; TAB and Unicode within descriptions remain intact.
- All numerical lexical tokens, including tiny E-values, remain text. No float64 score/E-value recomputation. Raw domain coordinates stay1-based inclusive with program-specific query/target roles. Reported and included domain counts each must not exceed observed domains; they are not ordered against one another.
- Complete JSON includes every hit/comment and UTF-8 line/description byte spans; original.txt preserves BOM and LF/CRLF. CSV protects formula-like text with a marked leading apostrophe; raw JSON/source values remain unchanged.
- Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
MARC record-field exportRead raw ISO2709 records into complete ordered field/subfield JSON and long CSV, preserving Unicode, repeated tags, indicators and directory/value byte spans.
Export catalogue data without flattening repeated fields or guessing an encoding. Check structural and encoding diagnostics alongside the original file.
Steps
- Select one complete .mrc byte stream and choose its declared encoding profile.
- Inspect record/field counts, unverified overrides and record diagnostics.
- Download all ordered fields and original bytes; use value spans to locate data in original.mrc.
Available options
- Encoding profile
- Require UTF-8 in leader · Known UTF-8, blank leader: unverified override
Capabilities and limits
- One raw file up to10MiB,10,000 records,100,000 fields and200,000 subfields. Complete original.mrc+records.json+fields.csv must fit64MiB. Decimal framing permits fields up to9,999 bytes and records up to99,999 bytes; budgets are coupled.
- Declared profile:24-byte ASCII leader,12-byte directory entries, entry map4500, two indicators and two-byte subfield identifier. Numeric/alphabetic single-case three-character tags and local ASCII subfield symbols are retained. Outer BOM, gaps, overlaps, unindexed data and malformed framing reject.
- Leader a declares UTF-8. A blank leader encoding requires the explicit known-UTF-8 override and remains unverified. Other encoding schemes and MARC8 conversion are Unsupported. Invalid UTF-8 rejects; a literal BOM inside a field is retained.
- Directory order and physical data order remain separately recorded, with all repeated fields/subfields and original values. Known ID/005/order issues are diagnostics, not a claim of full MARC21 catalogue, tag-meaning or character-repertoire validity. No MARCXML or wide-CSV flattening.
- fields.csv contains every control/subfield value. Formula-like cells get a leading apostrophe and csvLiteralPrefixAdded=1; recover original values from JSON/source or the explicit flag. The UI table displays original values.
- Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
Multilocus sequence-matrix builderJoin already aligned DNA FASTA loci by exact taxon ID and export a complete FASTA/NEXUS supermatrix, locus partitions, presence map and original sources.
Select ordered aligned loci. Union explicitly fills absent taxa; intersection keeps only shared taxa. This tool concatenates existing alignments and creates charsets.
Steps
- Select the loci in the desired order; each file must already contain an alignment.
- Choose union/intersection and the missing-locus symbol, then check taxa, columns, filled cells and presence rows.
- Download complete concatenated.fasta, concatenated.nex, presence.csv, report.json and all original files.
Available options
- Taxon set
- Union; fill absent loci · Intersection only
- Absent-locus symbol
- ? · - · N
Capabilities and limits
- Up to20 files with10MiB total,5,000 output taxa,500,000 total columns,10,000,000 matrix cells,100,000 source records and300,000 physical lines. Each line≤65,536 bytes and header≤4,096 bytes. Complete derived files and originals≤48MiB.
- Limits apply together:5,000 taxa ×500,000 columns exceeds the10million-cell guard. The source-record maximum follows20 loci ×5,000 unique source taxa; it is not an additional independent100,001-record supported boundary.
- UTF-8 with optional BOM, LF/CRLF, aligned DNA IUPAC symbols plus ? and -. Sequences must be equally long within each locus; IDs are the case-sensitive first header token and unique within a locus. Full headers, source spans, case and original bytes stay intact.
- RNA/protein, comments and embedded sequence spaces are Unsupported. No aligner, fuzzy species-name matching or phylogenetic inference. Output taxon order follows first occurrence in source-file/record order; file order defines locus columns.
- NEXUS exports DNA matrix and fixed locus_ordinal charsets with complete original file labels in JSON. Every taxon/locus presence or filled state appears in presence.csv/report.json. Formula-like CSV taxon labels are protected with a leading apostrophe; unchanged labels remain in JSON.
- Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.