PDB declared chain sequence exporter
Export legacy PDB protein SEQRES declarations as chain-labelled FASTA, with fixed-column residue provenance, explicit MODRES mappings and unknown positions.
- 1Add input
- 2Adjust settings
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Extract the sequence declared for a selected chain. SEQRES declarations can differ from residues with coordinates; the report keeps those meanings separate.
How to use this tool
- Open an ASCII PDB with SEQRES records and select all chains or one exact chain character.
- Compare input/selected counts and inspect unknown or MODRES-mapped residue positions.
- Download declared chain FASTA and complete source-position reports.
Supported inputs and limits
One ASCII legacy text file or paste up to10 MiB; 200,000 lines, 95 printable ASCII chain IDs including blank, and100,000 residues. Complete downloads up to20 MiB. Each chain numRes uses its four-digit field and must match all collected residues.
Fixed PDB columns only: serial starts1 and increments for each chain, count stays constant, up to13 populated three-character slots per SEQRES line. Invalid separators, gaps, short codes, extra residue fields or count mismatches reject. Nucleotide profiles, mmCIF and non-ASCII input are unsupported.
Maps20 standard amino acids plus ASX/GLX/UNK. Unique chain+residue-name MODRES parent mappings are used; conflicting parents reject. Other uppercase/digit three-character protein names becomeX with original name, parent, one-based position, source line and column. Unknowns remain visible.
Select all chains or one exact character, including a space for blank. A missing selected chain rejects. FASTA chain-uXXXX identifiers reversibly encode the original ASCII code; blank is chain-u0020. Selected reports still show complete input chain/residue counts. No SEQRES means no-declared-sequence; ATOM is never substituted.
FASTA, JSON and CSV include every selected residue. Table preview first 200 rows, cells up to2,000 characters; copying reports above20,000 characters gives a labelled preview. This converts declarations and does not validate observed coordinates, infer missing residues or establish biological facts.
Worked example
Example input
SEQRES 1 A 21 GLY ILE VAL GLU GLN CYS CYS THR SER ILE CYS SER LEU SEQRES 2 A 21 TYR GLN LEU GLU ASN TYR CYS ASN SEQRES 1 B 3 MSE UNK ALA MODRES 1ABC MSE B 1 MET SELENOMETHIONINE
Example options
{"secondary":"","params":{"chainMode":"all","chainId":"A","spreadsheetSafe":true}}Example output
{"format":"legacy-PDB3.3-protein-SEQRES","summary":{"inputLines":4,"inputChains":2,"inputResidues":24,"selectedChains":2,"selectedResidues":24,"unknownResidues":1,"status":"declared-sequences"},"chainEncoding":"FASTA identifier chain-uXXXX encodes the original printable ASCII code point in uppercase hexadecimal; blank is chain-u0020.","chains":[{"chain":"A","chainCode":65,"declaredCount":21,"lastSerial":2,"residues":[{"position":1,"name":"GLY","parent":"GLY","letter":"G","status":"standard","line":1,"column":20},{"position":2,"name":"ILE","parent":"ILE","letter":"I","status":"standard","line":1,"column":24},{"position":3,"name":"VAL","parent":"VAL","letter":"V","status":"standard","line":1,"column":28},{"position":4,"name":"GLU","parent":"GLU","letter":"E","status":"standard","line":1,"column":32},{"position":5,"name":"GLN","parent":"GLN","letter":"Q","status":"standard","line":1,"column":36},{"position":6,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":40},{"position":7,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":44},{"position":8,"name":"THR","parent":"THR","letter":"T","status":"standard","line":1,"column":48},{"position":9,"name":"SER","parent":"SER","letter":"S","status":"standard","line":1,"column":52},{"position":10,"name":"ILE","parent":"ILE","letter":"I","status":"standard","line":1,"column":56},{"position":11,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":60},{"position":12,"name":"SER","parent":"SER","letter":"S","status":"standard","line":1,"column":64},{"position":13,"name":"LEU","parent":"LEU","letter":"L","status":"standard","line":1,"column":68},{"position":14,"name":"TYR","parent":"TYR","letter":"Y","status":"standard","line":2,"column":20},{"position":15,"name":"GLN","parent":"GLN","letter":"Q","status":"standard","line":2,"column":24},{"position":16,"name":"LEU","parent":"LEU","letter":"L","status":"standard","line":2,"column":28},{"position":17,"name":"GLU","parent":"GLU","letter":"E","status":"standard","line":2,"column":32},{"position":18,"name":"ASN","parent":"ASN","letter":"N","status":"standard","line":2,"column":36},{"position":19,"name":"TYR","parent":"TYR","letter":"Y","status":"standard","line":2,"column":40},{"position":20,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":2,"column":44},{"position":21,"name":"ASN","parent":"ASN","letter":"N","status":"standard","line":2,"column":48}],"sequence":"GIVEQCCTSICSLYQLENYCN"},{"chain":"B","chainCode":66,"declaredCount":3,"lastSerial":1,"residues":[{"position":1,"name":"MSE","parent":"MET","letter":"M","status":"MODRES-mapped","line":3,"column":20},{"position":2,"name":"UNK","parent":"UNK","letter":"X","status":"unknown-residue","line":3,"column":24},{"position":3,"name":"ALA","parent":"ALA","letter":"A","status":"standard","line":3,"column":28}],"sequence":"MXA"}],"modres":[{"line":4,"chain":"B","name":"MSE","parent":"MET"}],"scope":"Declared SEQRES only, no ATOM fallback/observed-residue/biological validation. 20 standard amino acids plus ASX/GLX/UNK; unique MODRES chain+name mapping. Unknown three-character names become X with source positions. Nucleotides/mmCIF unsupported."}When something does not work
Check ASCII legacy PDB fixed columns, three-character protein residue slots, sequential serials, constant matching numRes and unambiguous MODRES parents. Select an existing chain, or re-export unsupported mmCIF/nucleotide data with its producer.
Frequently asked questions
Why can this differ from the coordinate sequence?
SEQRES states the declared sequence. ATOM/HETATM describe coordinate observations and can omit residues. This tool never fills one from the other.
How is a blank chain represented?
Its original character isASCII32. Select a single space; FASTA uses the reversible identifierchain-u0020.
Does every modified residue become a known amino acid?
No. Only a unique explicit MODRES parent or the stated standard mapping is used. Unknown names remainX with their source positions; conflicting mappings reject.
Documentation & further reading
Related tools
JSON formatting workspace
Format or minify strict JSON, sort object keys, and encode or decode strings while preserving raw number tokens.
Regex tester
Try a pattern and see what it matches in your text.
Compare text
See what changed, side by side.
HTML formatter
Format HTML indentation so its structure is easier to read.