Identical peptides can have different source positions
Understand why missed-cleavage enumeration needs occurrences, explicit cut rules and source evidence rather than a deduplicated list of sequences.
A maximum depth includes shorter windows
Klemens Fröhlich’s question starts from GRGKA and asks for possible peptides when cuts may be missed. The author’s later comments clarify that the list should contain all combinations up to the chosen depth. That makes a sliding window over consecutive cleavage segments useful; skipping only alternating cuts would miss valid windows.
With author-KR, GRGKA has three unmissed segments: GR, GK and A. A maximum of one missed cleavage retains those three and adds GRGK and GKA. A maximum of two additionally includes GRGKA. The five-row depth-one result is complete even though there are only two newly joined windows.
A sequence is not the same identity as an occurrence
A set of peptide strings answers which distinct sequences exist. An occurrence table answers where each window came from. Those questions differ when a protein repeats a residue or motif. Under author-KR at depth zero, KRK has K at [0,1), R at [1,2) and another K at [2,3).
Both K rows carry evidence. Keeping the coordinates and actual missed count allows downstream work to group by sequence deliberately while retaining every source interval. Deduplicating during enumeration would make the original positions impossible to reconstruct from the remaining row.
| start0 | end0 | missedCleavages | sequence |
|---|---|---|---|
| 0 | 1 | 0 | K |
| 1 | 2 | 0 | R |
| 2 | 3 | 0 | K |
The cut rule belongs with the result
The original author’s K/R split cuts before P. Expasy trypsin normally does not, while WKP and MRP are explicit exceptions. For AKP, the author rule produces AK and P at depth zero; Expasy produces the single AKP window. A label saying only trypsin would hide that consequential choice.
The report therefore keeps an explicit profile and the complete cuts0 array. Coordinates are zero-based and half-open: the end boundary is excluded. A terminal K or R does not produce an empty trailing peptide; the sequence end remains one boundary.
Neither rule estimates what an experiment will observe. Window enumeration does not calculate peptide masses, modifications, concentrations, yields or digestion probabilities. Select the rule for the intended theoretical analysis, then keep that choice visible.
Completeness needs a source and a visible refusal
One FASTA record keeps coordinates attached to one protein. Its full header, including spaces, remains in the report; the byte-exact original keeps BOM, blank lines and line endings. A second record is rejected as a whole so it cannot disappear while a first-record result looks complete.
Increasing missed depth grows the number of windows. The full count is checked before occurrence output allocation, with a protective ceiling of 200,000. At depth 10, 18,186 cleavage segments produce 199,991 occurrences; 18,187 produce 200,002 and are refused. The cap cannot be presented as an attainable exact 200,000-row sample.
The on-screen preview ends at 4,000 Unicode code points, while report copy and all CSV rows remain complete. A refusal, cancellation or the 10-second read/load/Worker deadline yields no partial or late result. Download the full report, CSV and original together when the analysis must be checked by someone else.
- Keep the complete cuts0 array, selected profile and maximum depth with every occurrence table.
- Retain identical sequences at different positions; group them only after preserving the source intervals.
- Compare source name, byte length and SHA256 with the intended original before handing off positions.
- Use the full report copy or downloads to review long results; a screen preview does not represent the entire enumeration.
References
- Klemens Fröhlich — missed-cleavage combinations
Complete captured question, both answers and all comments, including author clarifications about all depths. The record does not prove author completion, browser preference or market volume.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Tryptic peptide positionsEnumerate every theoretical peptide occurrence from one protein sequence or FASTA, up to a chosen missed-cleavage depth, with positions, CSV and exact source bytes.
Choose a cleavage rule and maximum missed-cleavage depth for one protein. Keep every peptide occurrence, including identical sequences at different positions, and export the complete report, CSV and original.
Steps
- Choose sequence or single-record FASTA, a cleavage rule and an integer maximum missed-cleavage depth.
- Paste the protein or select its original UTF-8 file.
- Enumerate and inspect the source length, cut boundaries and complete occurrence count.
- Copy the full JSON report or download the CSV, report and original; retain the positions when comparing repeated peptide sequences.
Available options
- Input format
- Protein sequence · Single-record FASTA
- Cleavage rule
- Author K/R (including before P) · Expasy trypsin (WKP/MRP exceptions)
- Maximum missed cleavages
- 1
Integer 0–10; includes every depth from zero to this maximum.
Capabilities and limits
- Paste one protein or select one UTF-8 file, up to 1 MiB. A selected file takes priority over pasted text; clear the selection to use the text. Its nonempty name must be scalar Unicode, at most 512 UTF-8 bytes, with no control characters or BOM. Choose the format explicitly; the extension does not determine it.
- Residues must use only the uppercase ASCII letters ACDEFGHIKLMNPQRSTVWY, with a nonempty total of at most 20,000 residues. Lowercase letters, spaces, gaps, ambiguous residues, stop symbols and modifications are rejected rather than normalized.
- Sequence format accepts one physical sequence line, with an optional final LF or CRLF. FASTA accepts exactly one record: the first line starts with > and has a nonempty full header of at most 65,536 UTF-8 bytes, excluding >. Header tabs, controls and BOM are rejected. Blank lines after the header are allowed; each nonempty sequence line must contain only supported residues. A second record rejects the whole input.
- One leading UTF-8 BOM is accepted and recorded. LF, CRLF and their mixture are preserved in the original; bare CR is rejected. At most 50,000 physical lines, counting the FASTA header and blank lines. No header is reduced to its first token.
- Author K/R cuts after every K or R, including before P. Expasy trypsin follows the Pyteomics 4.7.5 rule: normally no cut before P, with explicit WKP and MRP exceptions. Both retain the terminal boundary. These are theoretical rules, not a prediction of experimental digestion.
- The maximum missed-cleavage depth must be an integer from 0 through 10. Every consecutive cleavage-segment window with an actual missed count from 0 through that maximum is included. Records are ordered by end0, then start0; identical peptide sequences at different positions remain separate.
- start0 is included and end0 is excluded in zero-based residue coordinates. cuts0 contains all boundaries, including 0 and the sequence length. The complete occurrence count is checked before occurrence output is allocated; the protective limit is 200,000 occurrences. The largest naturally admissible count is 199,991; the first crossing at depth 10 is 200,002.
- Download peptides.csv, result-report.json and source-protein.txt or source-protein.fasta. CSV includes every start0, end0, missedCleavages and sequence. The report keeps the full header, source identity, all cuts, all occurrences and export metadata. The original is byte-exact. All three downloads together are limited to 32 MiB; no partial result is returned.
- The screen previews at most 4,000 Unicode code points; copy retains the complete JSON report. The 10-second deadline covers file reading, loading and Worker processing. Cancellation or timeout terminates unfinished processing and publishes no late or partial result. Contents and filenames stay in this browser.