Neatbo.

PDB declared chain sequence exporter

Export legacy PDB protein SEQRES declarations as chain-labelled FASTA, with fixed-column residue provenance, explicit MODRES mappings and unknown positions.

Browser-local processingInputASCII legacy PDB3.3 SEQRESOutputDeclared FASTA / residue JSON / CSVUp to 10 MiB per file · File limit: 1
  1. 1Add input
  2. 2Adjust settings
  3. 3Get your result

Tool input and files are processed in this browser without being uploaded.

Your input

Inputs are kept temporarily in this tab when switching tools. Refreshing or closing clears them; large results may need to be regenerated.

⌘ / Ctrl + Enter to run

or drag and drop it here

Files stay on this device. Your originals stay unchanged.

.pdb · .ent · .txt

Up to 10 MiB per file · File limit: 1

    0 characters · 0 bytes
    Options

    Complete the required options first. You can keep the defaults for the rest.

    Preparing the tool…

    Before you start

    Extract the sequence declared for a selected chain. SEQRES declarations can differ from residues with coordinates; the report keeps those meanings separate.

    How to use this tool

    1. Open an ASCII PDB with SEQRES records and select all chains or one exact chain character.
    2. Compare input/selected counts and inspect unknown or MODRES-mapped residue positions.
    3. Download declared chain FASTA and complete source-position reports.

    Supported inputs and limits

    One ASCII legacy text file or paste up to10 MiB; 200,000 lines, 95 printable ASCII chain IDs including blank, and100,000 residues. Complete downloads up to20 MiB. Each chain numRes uses its four-digit field and must match all collected residues.

    Fixed PDB columns only: serial starts1 and increments for each chain, count stays constant, up to13 populated three-character slots per SEQRES line. Invalid separators, gaps, short codes, extra residue fields or count mismatches reject. Nucleotide profiles, mmCIF and non-ASCII input are unsupported.

    Maps20 standard amino acids plus ASX/GLX/UNK. Unique chain+residue-name MODRES parent mappings are used; conflicting parents reject. Other uppercase/digit three-character protein names becomeX with original name, parent, one-based position, source line and column. Unknowns remain visible.

    Select all chains or one exact character, including a space for blank. A missing selected chain rejects. FASTA chain-uXXXX identifiers reversibly encode the original ASCII code; blank is chain-u0020. Selected reports still show complete input chain/residue counts. No SEQRES means no-declared-sequence; ATOM is never substituted.

    FASTA, JSON and CSV include every selected residue. Table preview first 200 rows, cells up to2,000 characters; copying reports above20,000 characters gives a labelled preview. This converts declarations and does not validate observed coordinates, infer missing residues or establish biological facts.

    Worked example

    Example input

    SEQRES   1 A   21  GLY ILE VAL GLU GLN CYS CYS THR SER ILE CYS SER LEU
    SEQRES   2 A   21  TYR GLN LEU GLU ASN TYR CYS ASN
    SEQRES   1 B    3  MSE UNK ALA
    MODRES 1ABC MSE B    1  MET  SELENOMETHIONINE
    
    Example options
    {"secondary":"","params":{"chainMode":"all","chainId":"A","spreadsheetSafe":true}}

    Example output

    {"format":"legacy-PDB3.3-protein-SEQRES","summary":{"inputLines":4,"inputChains":2,"inputResidues":24,"selectedChains":2,"selectedResidues":24,"unknownResidues":1,"status":"declared-sequences"},"chainEncoding":"FASTA identifier chain-uXXXX encodes the original printable ASCII code point in uppercase hexadecimal; blank is chain-u0020.","chains":[{"chain":"A","chainCode":65,"declaredCount":21,"lastSerial":2,"residues":[{"position":1,"name":"GLY","parent":"GLY","letter":"G","status":"standard","line":1,"column":20},{"position":2,"name":"ILE","parent":"ILE","letter":"I","status":"standard","line":1,"column":24},{"position":3,"name":"VAL","parent":"VAL","letter":"V","status":"standard","line":1,"column":28},{"position":4,"name":"GLU","parent":"GLU","letter":"E","status":"standard","line":1,"column":32},{"position":5,"name":"GLN","parent":"GLN","letter":"Q","status":"standard","line":1,"column":36},{"position":6,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":40},{"position":7,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":44},{"position":8,"name":"THR","parent":"THR","letter":"T","status":"standard","line":1,"column":48},{"position":9,"name":"SER","parent":"SER","letter":"S","status":"standard","line":1,"column":52},{"position":10,"name":"ILE","parent":"ILE","letter":"I","status":"standard","line":1,"column":56},{"position":11,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":1,"column":60},{"position":12,"name":"SER","parent":"SER","letter":"S","status":"standard","line":1,"column":64},{"position":13,"name":"LEU","parent":"LEU","letter":"L","status":"standard","line":1,"column":68},{"position":14,"name":"TYR","parent":"TYR","letter":"Y","status":"standard","line":2,"column":20},{"position":15,"name":"GLN","parent":"GLN","letter":"Q","status":"standard","line":2,"column":24},{"position":16,"name":"LEU","parent":"LEU","letter":"L","status":"standard","line":2,"column":28},{"position":17,"name":"GLU","parent":"GLU","letter":"E","status":"standard","line":2,"column":32},{"position":18,"name":"ASN","parent":"ASN","letter":"N","status":"standard","line":2,"column":36},{"position":19,"name":"TYR","parent":"TYR","letter":"Y","status":"standard","line":2,"column":40},{"position":20,"name":"CYS","parent":"CYS","letter":"C","status":"standard","line":2,"column":44},{"position":21,"name":"ASN","parent":"ASN","letter":"N","status":"standard","line":2,"column":48}],"sequence":"GIVEQCCTSICSLYQLENYCN"},{"chain":"B","chainCode":66,"declaredCount":3,"lastSerial":1,"residues":[{"position":1,"name":"MSE","parent":"MET","letter":"M","status":"MODRES-mapped","line":3,"column":20},{"position":2,"name":"UNK","parent":"UNK","letter":"X","status":"unknown-residue","line":3,"column":24},{"position":3,"name":"ALA","parent":"ALA","letter":"A","status":"standard","line":3,"column":28}],"sequence":"MXA"}],"modres":[{"line":4,"chain":"B","name":"MSE","parent":"MET"}],"scope":"Declared SEQRES only, no ATOM fallback/observed-residue/biological validation. 20 standard amino acids plus ASX/GLX/UNK; unique MODRES chain+name mapping. Unknown three-character names become X with source positions. Nucleotides/mmCIF unsupported."}

    When something does not work

    Check ASCII legacy PDB fixed columns, three-character protein residue slots, sequential serials, constant matching numRes and unambiguous MODRES parents. Select an existing chain, or re-export unsupported mmCIF/nucleotide data with its producer.

    Frequently asked questions

    Why can this differ from the coordinate sequence?

    SEQRES states the declared sequence. ATOM/HETATM describe coordinate observations and can omit residues. This tool never fills one from the other.

    How is a blank chain represented?

    Its original character isASCII32. Select a single space; FASTA uses the reversible identifierchain-u0020.

    Does every modified residue become a known amino acid?

    No. Only a unique explicit MODRES parent or the stated standard mapping is used. Unknown names remainX with their source positions; conflicting mappings reject.

    Documentation & further reading

    Related tools