Enumerate all peptide occurrences up to a missed-cleavage depth
Prepare one strict protein sequence or FASTA, choose the cleavage rule, and keep complete positions, CSV, report and original bytes.
Choose one protein and its actual format
Choose sequence for one physical residue line, or FASTA for one record whose first line starts with > and contains a nonempty header. The format choice is explicit; a filename extension does not change it. A selected file takes priority over pasted text. Clear the selection to use the text box.
Use only uppercase ASCII ACDEFGHIKLMNPQRSTVWY and at most 20,000 residues. Spaces, lowercase letters, ambiguous residues, gaps, stop symbols and modification syntax are rejected. FASTA may wrap its sequence across nonempty residue lines and may have blank lines after its header. A second header rejects the whole input instead of selecting the first protein.
UTF-8 input accepts one leading BOM, LF and CRLF, including mixed line endings; the original preserves them. Bare CR and malformed UTF-8 are rejected. The full FASTA header is limited to 65,536 UTF-8 bytes, excluding >, with no tabs, controls or BOM. Count its header and blank lines within the 50,000 physical-line limit. The file must be at most 1 MiB and its name at most 512 UTF-8 bytes.
Set the rule and the maximum depth
Author K/R cuts after every K or R, even before P. Expasy trypsin uses the Pyteomics 4.7.5 rule: it normally avoids K/R cuts before P, but WKP and MRP are explicit exceptions. Both include the terminal boundary. Keep the selected profile with the result so another analysis can reproduce the same cut positions.
Enter an integer maximum from 0 through 10. A maximum of 1 includes both unmissed single segments and adjacent two-segment windows; it does not request only the latter. Increasing the maximum adds longer consecutive windows up to that depth.
| Maximum depth | Included peptide windows | Occurrence count |
|---|---|---|
| 0 | GR; GK; A | 3 |
| 1 | GR; GRGK; GK; GKA; A | 5 |
| 2 | GR; GRGK; GK; GRGKA; GKA; A | 6 |
Read coordinates as occurrences
start0 is included and end0 is excluded in zero-based residue coordinates. For GRGKA, [0,2) reads GR, while [2,5) reads GKA. missedCleavages is the actual number of internal cut boundaries skipped by that particular window, not the selected maximum.
cuts0 includes 0, every selected-rule cut and the sequence length. Occurrences are sorted by end0 and then start0. Keep rows with identical sequences at different coordinates: KRK at depth 0 has K at [0,1) and at [2,3). There is no unique-sequence deduplication.
start0,end0,missedCleavages,sequence
0,2,0,GR
0,4,1,GRGK
2,4,0,GK
2,5,1,GKA
4,5,0,A
Keep all three complete deliveries
The screen previews at most 4,000 Unicode code points. Copy obtains the complete result-report.json, including all cuts and every occurrence. Download all three files when handing off the analysis; the original identifies the precise source behind those positions.
The report retains the full FASTA header or null for sequence input, the explicit format and rule, maximum depth, source name, byte length, SHA256, BOM and line-ending metadata, all counts, all cuts, all occurrences and export metadata. CSV rows retain the four occurrence fields without truncating the sequence. Original download filenames are fixed; the original source name stays in the report.
| File | What to check |
|---|---|
| peptides.csv | Every start0, end0, missedCleavages and complete sequence, in report order |
| result-report.json | Source identity, full header, profile, cuts0, all occurrences, counts and export sizes |
| source-protein.txt / source-protein.fasta | Exact effective source bytes, including BOM, header, blank lines and line endings |
- Before running, confirm the effective source, explicit format, cleavage rule and maximum depth. A selected file takes priority over pasted text.
- After completion, download all three files and compare the CSV data-row count, excluding the header, with occurrenceCount in result-report.json.
- Compare CSV rows with report occurrences in order, retaining identical sequences at different coordinates. Keep the original file with the report’s source identity.
Recover without a partial enumeration
The complete occurrence count is calculated before allocating the occurrence output. The protective cap is 200,000; admissible natural counts reach 199,991 and the first crossing at depth 10 is 200,002. There is no grammar-valid exact 200,000 case under the other bounds.
The 1 MiB input and combined 32 MiB download limits are protective ceilings. The full accepted source grammar is bounded by 185,540 bytes, and the complete report, CSV and original have a conservative combined bound of 20,141,739 bytes. Those stricter bounds prevent valid inputs or outputs from reaching the exact larger ceilings; they are not measured claims about browser memory.
The 10-second deadline starts before file reading and loading and covers Worker readiness and computation. Cancellation or timeout terminates unfinished processing; no partial files, report copy or late result are published. Keep the same selected File and parameters to retry. Correct a grammar error or reduce the protein or depth after a count refusal.
Processing stays in the browser without uploading input content or names. These outputs enumerate theoretical windows. They do not establish peptide mass, modifications, experimental yield or actual digestion probability.
References
- Klemens Fröhlich — missed-cleavage combinations
The captured complete question, both answers and all comments establish a protein/FASTA task and all windows up to the selected depth. No browser preference or completed-author result is established.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Tryptic peptide positionsEnumerate every theoretical peptide occurrence from one protein sequence or FASTA, up to a chosen missed-cleavage depth, with positions, CSV and exact source bytes.
Choose a cleavage rule and maximum missed-cleavage depth for one protein. Keep every peptide occurrence, including identical sequences at different positions, and export the complete report, CSV and original.
Steps
- Choose sequence or single-record FASTA, a cleavage rule and an integer maximum missed-cleavage depth.
- Paste the protein or select its original UTF-8 file.
- Enumerate and inspect the source length, cut boundaries and complete occurrence count.
- Copy the full JSON report or download the CSV, report and original; retain the positions when comparing repeated peptide sequences.
Available options
- Input format
- Protein sequence · Single-record FASTA
- Cleavage rule
- Author K/R (including before P) · Expasy trypsin (WKP/MRP exceptions)
- Maximum missed cleavages
- 1
Integer 0–10; includes every depth from zero to this maximum.
Capabilities and limits
- Paste one protein or select one UTF-8 file, up to 1 MiB. A selected file takes priority over pasted text; clear the selection to use the text. Its nonempty name must be scalar Unicode, at most 512 UTF-8 bytes, with no control characters or BOM. Choose the format explicitly; the extension does not determine it.
- Residues must use only the uppercase ASCII letters ACDEFGHIKLMNPQRSTVWY, with a nonempty total of at most 20,000 residues. Lowercase letters, spaces, gaps, ambiguous residues, stop symbols and modifications are rejected rather than normalized.
- Sequence format accepts one physical sequence line, with an optional final LF or CRLF. FASTA accepts exactly one record: the first line starts with > and has a nonempty full header of at most 65,536 UTF-8 bytes, excluding >. Header tabs, controls and BOM are rejected. Blank lines after the header are allowed; each nonempty sequence line must contain only supported residues. A second record rejects the whole input.
- One leading UTF-8 BOM is accepted and recorded. LF, CRLF and their mixture are preserved in the original; bare CR is rejected. At most 50,000 physical lines, counting the FASTA header and blank lines. No header is reduced to its first token.
- Author K/R cuts after every K or R, including before P. Expasy trypsin follows the Pyteomics 4.7.5 rule: normally no cut before P, with explicit WKP and MRP exceptions. Both retain the terminal boundary. These are theoretical rules, not a prediction of experimental digestion.
- The maximum missed-cleavage depth must be an integer from 0 through 10. Every consecutive cleavage-segment window with an actual missed count from 0 through that maximum is included. Records are ordered by end0, then start0; identical peptide sequences at different positions remain separate.
- start0 is included and end0 is excluded in zero-based residue coordinates. cuts0 contains all boundaries, including 0 and the sequence length. The complete occurrence count is checked before occurrence output is allocated; the protective limit is 200,000 occurrences. The largest naturally admissible count is 199,991; the first crossing at depth 10 is 200,002.
- Download peptides.csv, result-report.json and source-protein.txt or source-protein.fasta. CSV includes every start0, end0, missedCleavages and sequence. The report keeps the full header, source identity, all cuts, all occurrences and export metadata. The original is byte-exact. All three downloads together are limited to 32 MiB; no partial result is returned.
- The screen previews at most 4,000 Unicode code points; copy retains the complete JSON report. The 10-second deadline covers file reading, loading and Worker processing. Cancellation or timeout terminates unfinished processing and publishes no late or partial result. Contents and filenames stay in this browser.