Neatbo.

Tryptic peptide positions

Enumerate every theoretical peptide occurrence from one protein sequence or FASTA, up to a chosen missed-cleavage depth, with positions, CSV and exact source bytes.

Browser-local processingInputProtein sequence / single-record FASTAOutputComplete CSV, JSON report and originalUp to 1 MiB per file · File limit: 1
  1. 1Add input
  2. 2Adjust settings
  3. 3Get your result

Tool input and files are processed in this browser without being uploaded.

Your input

Inputs are kept temporarily in this tab when switching tools. Refreshing or closing clears them; large results may need to be regenerated.

⌘ / Ctrl + Enter to run

or drag and drop it here

Files stay on this device. Your originals stay unchanged.

.fasta · .fa · .faa · .fas · .txt

Up to 1 MiB per file · File limit: 1

    0 characters · 0 bytes
    Options

    Complete the required options first. You can keep the defaults for the rest.

    Integer 0–10; includes every depth from zero to this maximum.

    Preparing the tool…

    Before you start

    Choose a cleavage rule and maximum missed-cleavage depth for one protein. Keep every peptide occurrence, including identical sequences at different positions, and export the complete report, CSV and original.

    How to use this tool

    1. Choose sequence or single-record FASTA, a cleavage rule and an integer maximum missed-cleavage depth.
    2. Paste the protein or select its original UTF-8 file.
    3. Enumerate and inspect the source length, cut boundaries and complete occurrence count.
    4. Copy the full JSON report or download the CSV, report and original; retain the positions when comparing repeated peptide sequences.

    Supported inputs and limits

    Paste one protein or select one UTF-8 file, up to 1 MiB. A selected file takes priority over pasted text; clear the selection to use the text. Its nonempty name must be scalar Unicode, at most 512 UTF-8 bytes, with no control characters or BOM. Choose the format explicitly; the extension does not determine it.

    Residues must use only the uppercase ASCII letters ACDEFGHIKLMNPQRSTVWY, with a nonempty total of at most 20,000 residues. Lowercase letters, spaces, gaps, ambiguous residues, stop symbols and modifications are rejected rather than normalized.

    Sequence format accepts one physical sequence line, with an optional final LF or CRLF. FASTA accepts exactly one record: the first line starts with > and has a nonempty full header of at most 65,536 UTF-8 bytes, excluding >. Header tabs, controls and BOM are rejected. Blank lines after the header are allowed; each nonempty sequence line must contain only supported residues. A second record rejects the whole input.

    One leading UTF-8 BOM is accepted and recorded. LF, CRLF and their mixture are preserved in the original; bare CR is rejected. At most 50,000 physical lines, counting the FASTA header and blank lines. No header is reduced to its first token.

    Author K/R cuts after every K or R, including before P. Expasy trypsin follows the Pyteomics 4.7.5 rule: normally no cut before P, with explicit WKP and MRP exceptions. Both retain the terminal boundary. These are theoretical rules, not a prediction of experimental digestion.

    The maximum missed-cleavage depth must be an integer from 0 through 10. Every consecutive cleavage-segment window with an actual missed count from 0 through that maximum is included. Records are ordered by end0, then start0; identical peptide sequences at different positions remain separate.

    start0 is included and end0 is excluded in zero-based residue coordinates. cuts0 contains all boundaries, including 0 and the sequence length. The complete occurrence count is checked before occurrence output is allocated; the protective limit is 200,000 occurrences. The largest naturally admissible count is 199,991; the first crossing at depth 10 is 200,002.

    Download peptides.csv, result-report.json and source-protein.txt or source-protein.fasta. CSV includes every start0, end0, missedCleavages and sequence. The report keeps the full header, source identity, all cuts, all occurrences and export metadata. The original is byte-exact. All three downloads together are limited to 32 MiB; no partial result is returned.

    The screen previews at most 4,000 Unicode code points; copy retains the complete JSON report. The 10-second deadline covers file reading, loading and Worker processing. Cancellation or timeout terminates unfinished processing and publishes no late or partial result. Contents and filenames stay in this browser.

    Worked example

    Example input

    GRGKA
    Example options
    {"format":"sequence","profile":"author-KR","missedCleavages":1}

    Example output

    start0,end0,missedCleavages,sequence
    0,2,0,GR
    0,4,1,GRGK
    2,4,0,GK
    2,5,1,GKA
    4,5,0,A
    

    When something does not work

    Keep the original. Correct the reported encoding, alphabet, header, record count or parameter error before retrying. Reduce the protein or missed-cleavage depth if its full occurrence count exceeds the limit. After cancelling or timing out, rerun with the same selected file and parameters. Clear the file selection to process pasted text.

    Frequently asked questions

    Does depth 1 return only peptides with one missed cleavage?

    It returns both zero and one missed cleavage. For GRGKA under author-KR, depth 1 yields five occurrences: GR, GRGK, GK, GKA and A. Depth 2 also includes the complete GRGKA window.

    Why can the same peptide appear several times?

    Each row represents an occurrence in the protein. KRK at depth 0 under author-KR contains K at [0,1) and K at [2,3). Removing one would discard a source position.

    Why do the two cleavage profiles disagree before P?

    The author rule splits AKP into AK and P. Expasy keeps AKP together, while WKP and MRP are explicit exceptions that still split before P. Select the rule required for your analysis and keep its name in the report.

    Can I submit a FASTA containing several proteins?

    No. A second header rejects the entire input; no protein is silently dropped. Submit one record at a time. This tool does not compute masses, modifications, peptide yields or observed digestion probabilities.

    Documentation & further reading

    Related tools