GenBank CDS declared-protein export
Select CDS features by a literal qualifier value and export their unique declared translations, retaining all feature statuses, qualifiers and original byte spans.
- 1Add input
- 2Adjust settings
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Read existing protein declarations from a complete GenBank file. Missing, empty and repeated translations remain visible; the tool does not recompute proteins from DNA.
How to use this tool
- Select a complete GenBank file or paste UTF-8 text. Set the qualifier key, literal value and contains/equals choice.
- Inspect matched and exported counts, then check each unavailable CDS status instead of assuming every match has a protein.
- Download proteins.faa with the complete report and original.gb; preserve the source when handing off declared locations.
Supported inputs and limits
One UTF-8 nucleic GenBank, optional BOM and LF/CRLF, up to10MiB. Up to10,000 records,100,000 features,20,000 matched CDS,8,000,000 declared translation ASCII characters across all features and300,000 physical lines. Complete original, JSON, CSV and FASTA together must fit40MiB.
Physical lines and declared locations each allow65,536 UTF-16 units; up to100,000 qualifiers per feature; each ordinary decoded qualifier up to1,048,576 UTF-16 units. Selection key is an ASCII identifier up to128 units; literal selection up to4,096. Limits apply together; input and output caps constrain simultaneous maxima.
Recognizes the nucleic LOCUS prefix, exact fixed ASCII FEATURES columns and complete ORIGIN/CONTIG plus // structure. It does not validate full INSDC headers, nucleotide content/length, locations, codons, translation tables, biology or external references. Protein LOCUS, lone CR, invalid UTF-8, unsupported indentation and translation characters are rejected.
Contains/equals selection is literal and case-sensitive, without regular expressions. Missing qualifier and no match are distinct. Only exactly one nonempty translation declaration exports; no declaration is missing, one empty declaration is empty, and any repeated declaration is ambiguous even if one is empty.
Full JSON preserves every parsed record/feature, ordered repeated qualifiers and flags, quote form, declared location and exact UTF-8 source spans. FASTA headers reversibly include record/feature indices, LOCUS, all ID values and location. Only whitespace is removed from existing translation values; no biological validity is inferred.
CSV includes every CDS status and protects formula-like text with a leading apostrophe; JSON keeps the underlying value. Preview shows first200 rows and first2,000 UTF-16 units per long cell. Above20,000 units, copy is a report preview; all downloads remain complete. Originals retain BOM and LF/CRLF byte for byte.
Worked example
Example input
LOCUS scaffold_10 12 bp DNA linear UNA 01-JAN-2000
DEFINITION Synthetic test record; declared translations are not recomputed.
ACCESSION SYNTH0001
VERSION SYNTH0001.1
KEYWORDS .
FEATURES Location/Qualifiers
CDS 13110..15500
/ID="id"
/NRPS_PKS="Domain: PKS_KR"
/translation="MKL*"
/NRPS_PKS="type: other"
CDS join(3252..3263,3321..3443,3493..4428)
/ID="id"
/NRPS_PKS="Domain: PKS_KR"
/translation="MKL*"
/ID="CDS_13111"
ORIGIN
1 acgtacgtacgt
//
Example options
{"secondary":"","params":{"qualifierKey":"NRPS_PKS","literal":"PKS_","mode":"contains"}}Example output
{"profile":"UTF8 nucleic GenBank LOCUS/FEATURES and declared CDS translations; not full INSDC/schema/sequence/location/biological validation","selection":{"qualifierKey":"NRPS_PKS","literal":"PKS_","mode":"contains","valueMatching":"literal-case-sensitive"},"source":{"encoding":"UTF-8","bytes":707,"bomRetained":false},"records":[{"index":0,"locus":"scaffold_10","declaredLengthToken":"12","accession":["SYNTH0001"],"version":"SYNTH0001.1","featureIndexes":[0,1],"source":{"utf8Start":0,"utf8Bytes":707}}],"features":[{"index":0,"record":0,"key":"CDS","qualifiers":[{"key":"ID","quoted":true,"flag":false,"source":{"utf8Start":294,"utf8Bytes":30},"value":"id"},{"key":"NRPS_PKS","quoted":true,"flag":false,"source":{"utf8Start":324,"utf8Bytes":48},"value":"Domain: PKS_KR"},{"key":"translation","quoted":true,"flag":false,"source":{"utf8Start":372,"utf8Bytes":41},"value":"MKL*"},{"key":"NRPS_PKS","quoted":true,"flag":false,"source":{"utf8Start":413,"utf8Bytes":45},"value":"type: other"}],"source":{"utf8Start":260,"utf8Bytes":198},"declaredLocation":"13110..15500"},{"index":1,"record":0,"key":"CDS","qualifiers":[{"key":"ID","quoted":true,"flag":false,"source":{"utf8Start":518,"utf8Bytes":30},"value":"id"},{"key":"NRPS_PKS","quoted":true,"flag":false,"source":{"utf8Start":548,"utf8Bytes":48},"value":"Domain: PKS_KR"},{"key":"translation","quoted":true,"flag":false,"source":{"utf8Start":596,"utf8Bytes":41},"value":"MKL*"},{"key":"ID","quoted":true,"flag":false,"source":{"utf8Start":637,"utf8Bytes":37},"value":"CDS_13111"}],"source":{"utf8Start":458,"utf8Bytes":216},"declaredLocation":"join(3252..3263,3321..3443,3493..4428)"}],"cds":[{"featureIndex":0,"recordIndex":0,"locus":"scaffold_10","ids":["id"],"declaredLocation":"13110..15500","predicateQualifierIndexes":[1,3],"matchingQualifierIndexes":[1],"translationQualifierIndexes":[2],"status":"exported"},{"featureIndex":1,"recordIndex":0,"locus":"scaffold_10","ids":["id","CDS_13111"],"declaredLocation":"join(3252..3263,3321..3443,3493..4428)","predicateQualifierIndexes":[1],"matchingQualifierIndexes":[1],"translationQualifierIndexes":[2],"status":"exported"}],"proteinManifest":[{"featureIndex":0,"header":"r=0|f=0|locus=scaffold_10|cds=%5B%22id%22%5D|location=13110..15500","aminoAcids":4},{"featureIndex":1,"header":"r=0|f=1|locus=scaffold_10|cds=%5B%22id%22%2C%22CDS_13111%22%5D|location=join(3252..3263%2C3321..3443%2C3493..4428)","aminoAcids":4}],"summary":{"records":1,"features":2,"cds":2,"matchedCDS":2,"exportedProteins":2,"exportedAminoAcids":8,"declaredTranslationCharacters":8,"physicalLines":19,"unavailableMatches":0},"unreviewed":["ORIGIN/CONTIG nucleotide content and length","location interpretation/extraction","other record headers","codon_start/transl_table consistency","biological protein validity","external references"]}When something does not work
Supply a complete nucleic record with supported fixed columns and encoding. Fix the literal selection or duplicate declaration at its source; never delete evidence to guess a protein. Reduce input or complete output after a limit rejection. Cancellation and failure publish no partial files.
Frequently asked questions
Does this translate DNA?
No. It exports existing declared translation values only; DNA, codon settings and translation consistency remain unreviewed.
What if translation appears twice?
Any repeated declaration is ambiguous and is not exported, including one empty and one nonempty declaration. All declarations stay in the report.
Are locations interpreted?
No. join, complement, fuzzy and remote locations are retained as declared strings with source spans. They are not used to extract nucleotides.
Documentation & further reading
Related tools
JSON formatting workspace
Format or minify strict JSON, sort object keys, and encode or decode strings while preserving raw number tokens.
Regex tester
Try a pattern and see what it matches in your text.
Compare text
See what changed, side by side.
HTML formatter
Format HTML indentation so its structure is easier to read.