Neatbo.

Preserve scientific records and alignments

Interpret accession gaps, repeated catalogue fields and missing loci before consuming derived tables and sequence matrices.

A flat table can hide a lost declaration

An accession is not interchangeable with a display name. A repeated MARC field is not a duplicate to remove by default. A missing taxon/locus block is not an observed gap sequence. Check these distinctions before consuming derived data.

  • Match HMMER query and target accessions to the selected program before joining tables by identifier.
  • Keep MARC repeated fields and subfields in source order; carry any explicit UTF-8 override status with the export.
  • Choose union or intersection for the aligned loci, then check the presence map and partition coordinates before handing off the matrix.

Compare the handoff to the source

For HMMER, inspect every retained query/target accession and description, and match domain roles to the selected program. Full JSON keeps raw numerical tokens even when a CSV importer turns them into floating numbers. For MARC, compare ordered indicators/subfields and physical byte spans; UTF-8 override status must travel with exported data.

Interpret derived meaning
ObservationRetained meaningNext action
Tiny E-value tokenOriginal decimal spellingUse complete report, do not infer statistical recomputation
Blank MARC encoding with overrideKnown UTF-8 assertion remains unverifiedReview source encoding; do not call it a MARC8 conversion
Filled missing locusTaxon absent in this source locusCheck presence map before downstream analysis

Preserve order and explain derived values

Supermatrix union keeps taxa absent from some loci by filling the declared symbol and recording presence; intersection deliberately removes taxa not shared by all loci. Both choices are part of the task and remain in the complete report. Existing case and full source headers remain intact.

A NEXUS charset points to concatenated columns, not to a newly inferred biological region. The source order and complete partition map let a recipient reverse that coordinate handoff. No alignment or phylogeny is computed.

Preview and independent checks have different scope

A 200-row panel is useful for checking the current choice, but the downloads are the full handoff. Formula-like CSV text is protected; JSON/original are the authority for unchanged values. Complete-output caps can reject an otherwise legal input, so reduce the actual combined task rather than accepting truncated results.

Mature search/catalogue/alignment readers check declared representations. They may group duplicate hits, normalize whitespace or shift coordinate origins; such differences are explained separately and never used to rewrite source bytes. Local Worker and implementation checks do not stand in for final production-browser acceptance.

References

  • HMMER export task

    First-person tblout-to-DataFrame task; query accession must not disappear.

  • MARC Unicode export task

    First-person Czech MARC-to-CSV encoding problem; original linked file was not obtained.

  • Multilocus workflow feedback

    Researcher describes missing taxa and charset preparation. Their1500-locus speed statement concerns other software.

Tools in this category

Expand a tool to see its steps, options and supported formats, then open its workspace.

HMMER result-table exportExport every HMMER3 protein-table hit with query and target accessions, original numeric tokens, free-text descriptions and exact source byte spans.

Choose the producing program and table type, then inspect complete hits without rerunning HMMER or changing statistical tokens.

Steps

  1. Choose tblout/domtblout and the actual producing program; paste or select one complete file.
  2. Check accessions, descriptions and domain coordinate roles; an empty hit table is distinct from a parse error.
  3. Download complete report.json, records.csv and original.txt.

Available options

Table type
tblout —18 fixed fields · domtblout —22 fixed fields
Producing program
hmmscan · hmmsearch · phmmer

Capabilities and limits

  • One10MiB UTF-8 protein HMMER3 tblout or domtblout;100,000 hit records;300,000 physical lines;65,536 bytes per physical line. Complete original+JSON+CSV must fit48MiB. Budgets apply together.
  • Explicit hmmscan/hmmsearch/phmmer and18/22-field table profiles. Header program/version contradictions, HMMER2/nhmmer, fixed-field TAB, Unicode-space identifiers and ambiguous hash-prefixed target rows are Unsupported. ASCII spaces separate fixed fields; TAB and Unicode within descriptions remain intact.
  • All numerical lexical tokens, including tiny E-values, remain text. No float64 score/E-value recomputation. Raw domain coordinates stay1-based inclusive with program-specific query/target roles. Reported and included domain counts each must not exceed observed domains; they are not ordered against one another.
  • Complete JSON includes every hit/comment and UTF-8 line/description byte spans; original.txt preserves BOM and LF/CRLF. CSV protects formula-like text with a marked leading apostrophe; raw JSON/source values remain unchanged.
  • Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
Open HMMER result-table export →
MARC record-field exportRead raw ISO2709 records into complete ordered field/subfield JSON and long CSV, preserving Unicode, repeated tags, indicators and directory/value byte spans.

Export catalogue data without flattening repeated fields or guessing an encoding. Check structural and encoding diagnostics alongside the original file.

Steps

  1. Select one complete .mrc byte stream and choose its declared encoding profile.
  2. Inspect record/field counts, unverified overrides and record diagnostics.
  3. Download all ordered fields and original bytes; use value spans to locate data in original.mrc.

Available options

Encoding profile
Require UTF-8 in leader · Known UTF-8, blank leader: unverified override

Capabilities and limits

  • One raw file up to10MiB,10,000 records,100,000 fields and200,000 subfields. Complete original.mrc+records.json+fields.csv must fit64MiB. Decimal framing permits fields up to9,999 bytes and records up to99,999 bytes; budgets are coupled.
  • Declared profile:24-byte ASCII leader,12-byte directory entries, entry map4500, two indicators and two-byte subfield identifier. Numeric/alphabetic single-case three-character tags and local ASCII subfield symbols are retained. Outer BOM, gaps, overlaps, unindexed data and malformed framing reject.
  • Leader a declares UTF-8. A blank leader encoding requires the explicit known-UTF-8 override and remains unverified. Other encoding schemes and MARC8 conversion are Unsupported. Invalid UTF-8 rejects; a literal BOM inside a field is retained.
  • Directory order and physical data order remain separately recorded, with all repeated fields/subfields and original values. Known ID/005/order issues are diagnostics, not a claim of full MARC21 catalogue, tag-meaning or character-repertoire validity. No MARCXML or wide-CSV flattening.
  • fields.csv contains every control/subfield value. Formula-like cells get a leading apostrophe and csvLiteralPrefixAdded=1; recover original values from JSON/source or the explicit flag. The UI table displays original values.
  • Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
Open MARC record-field export →
Multilocus sequence-matrix builderJoin already aligned DNA FASTA loci by exact taxon ID and export a complete FASTA/NEXUS supermatrix, locus partitions, presence map and original sources.

Select ordered aligned loci. Union explicitly fills absent taxa; intersection keeps only shared taxa. This tool concatenates existing alignments and creates charsets.

Steps

  1. Select the loci in the desired order; each file must already contain an alignment.
  2. Choose union/intersection and the missing-locus symbol, then check taxa, columns, filled cells and presence rows.
  3. Download complete concatenated.fasta, concatenated.nex, presence.csv, report.json and all original files.

Available options

Taxon set
Union; fill absent loci · Intersection only
Absent-locus symbol
? · - · N

Capabilities and limits

  • Up to20 files with10MiB total,5,000 output taxa,500,000 total columns,10,000,000 matrix cells,100,000 source records and300,000 physical lines. Each line≤65,536 bytes and header≤4,096 bytes. Complete derived files and originals≤48MiB.
  • Limits apply together:5,000 taxa ×500,000 columns exceeds the10million-cell guard. The source-record maximum follows20 loci ×5,000 unique source taxa; it is not an additional independent100,001-record supported boundary.
  • UTF-8 with optional BOM, LF/CRLF, aligned DNA IUPAC symbols plus ? and -. Sequences must be equally long within each locus; IDs are the case-sensitive first header token and unique within a locus. Full headers, source spans, case and original bytes stay intact.
  • RNA/protein, comments and embedded sequence spaces are Unsupported. No aligner, fuzzy species-name matching or phylogenetic inference. Output taxon order follows first occurrence in source-file/record order; file order defines locus columns.
  • NEXUS exports DNA matrix and fixed locus_ordinal charsets with complete original file labels in JSON. Every taxon/locus presence or filled state appears in presence.csv/report.json. Formula-like CSV taxon labels are protected with a leading apostrophe; unchanged labels remain in JSON.
  • Preview and copy contain at most200 rows/2,000 UTF-16 units per long table cell; above20,000 units the copied report is a compact preview. Full JSON, CSV and original downloads remain complete. All output limits are combined and enforced atomically.
Open Multilocus sequence-matrix builder →

Tools used in this article

HMMER result-table export →Export every HMMER3 protein-table hit with query and target accessions, original numeric tokens, free-text descriptions and exact source byte spans.MARC record-field export →Read raw ISO2709 records into complete ordered field/subfield JSON and long CSV, preserving Unicode, repeated tags, indicators and directory/value byte spans.Multilocus sequence-matrix builder →Join already aligned DNA FASTA loci by exact taxon ID and export a complete FASTA/NEXUS supermatrix, locus partitions, presence map and original sources.