Count BED overlaps with explicit reference and sample roles
Prepare compatible coordinate spans, preserve matrix column order, and verify the complete source package.
Assign roles before reading
Choose the reference first. Its logical interval rows become output rows in their original order. Select 1–20 samples and use Move up/Move down to arrange count columns. Filename extensions do not change these roles.
Use files from the same genome assembly and compatible chromosome spelling. The tool compares identifiers case-sensitively; chr1 and Chr1 do not match. It cannot detect an incorrect assembly label.
chr1 0 1000
chr1 1000 2000
Check the supported table
Each record needs 3–12 nonempty ASCII-space/TAB-delimited tokens. First three tokens are chromosome/start/end; extra tokens stay opaque. Comments must start with # in column zero. Whole original bytes, comments and line endings remain in sidecars.
Use nonempty half-open spans. A sample ending where the reference starts is adjacent and adds zero; duplicate and nested sample records add separately. A physically empty file is refused, while a nonempty comment-only sample contributes a zero column.
| Field or column | In this example | Meaning and check |
|---|---|---|
| Chromosome | chr1 | Match the original spelling and case; use compatible genome assemblies. |
| Start | 0 / 1000 | Zero-based and inclusive; must be a nonnegative decimal integer. |
| End | 1000 / 2000 | Exclusive and greater than start; touching endpoints do not overlap. |
| Extra reference fields | None here; columns 4–12 are allowed | Keep all tokens without interpreting strand or blocks; row widths may differ. |
| Appended sample counts | First row 2, 2; second row 1, 0 | Follow sample-001, then sample-002; these columns follow all reference tokens. |
chr1 20 50
chr1 30 40
chr1 1500 1900
Check the second sample and column order
Place sample-002.bed after sample-001.bed. Both of its records overlap reference [0,1000), and neither overlaps [1000,2000), so its count column is 2, then 0.
chr1 68 786
chr1 899 987
Retain the complete handoff
Copy returns the whole matrix, even when the screen preview stops at 4,000 code points. Download overlap-matrix.bed, result-report.json, source-reference.bed and every source-sample-NNN.bed. Compare all source byte lengths and SHA256 values with the intended inputs.
The report preserves every reference token and count, the ordered sample roles/names, interval totals, BOM/line-ending information and all export sizes/hashes. The matrix has no header, so retain the report alongside it.
All downloads jointly allow 32 MiB. Refusal, cancel or the 10-second host deadline creates no partial results. Correct a refused source or reduce the work, then rerun; cancellation retains the selected original Files and order.
chr1 0 1000 2 2
chr1 1000 2000 1 0
- Select reference.bed first, then sample-001.bed and sample-002.bed in that order. Check roles, genome assembly and chromosome spelling before running.
- For this example, check counts 2,2 and 1,0. Download the report and match each appended count column to its ordered sample source.
- If a file is refused, check start < end, UTF-8 and line endings, 3–12 fields, and the stated combined input and complete-download limits. Repair and reselect the source.
- After cancellation or timeout, rerun the retained files in the same order. Copy or download after the new run succeeds, and verify every original sidecar byte length and SHA256.
References
- Reference-by-sample overlap matrix request
Complete source question and both answers retained in the admission record. No author success, browser preference or demand-volume claim.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
BED overlap count matrixCount each sample interval that overlaps each reference span, including duplicate and nested records. Keep every reference field, the chosen sample order, exact originals and a complete source report.
Count each sample interval that overlaps each reference span, including duplicate and nested records. Keep every reference field, the chosen sample order, exact originals and a complete source report.
Steps
- Choose the reference file first, then the sample files.
- Move samples up/down until their order matches the intended matrix columns.
- Run and inspect reference/sample totals and complete download sizes.
- Copy the complete matrix or download the matrix, report and every original. Verify coordinate assembly compatibility in your analysis workflow.
Capabilities and limits
- Choose one reference BED file and 1–20 sample BED files in explicit column order. All original sources together allow 4 MiB and 100,000 interval records. A physically empty file is refused; a nonempty blank/comment-only sample contributes zeros. The reference must contain at least one interval.
- Each record has 3–12 nonempty tokens, separated by ASCII spaces or tabs. The first three are chromosome, start and end; every remaining token is opaque and retained. Leading/trailing ASCII separators normalize in the matrix. Quotes and backslashes are literal; empty fields or whitespace inside a field are unsupported.
- Chromosome and source labels allow 512 UTF-8 bytes; each coordinate or extra token allows 65,536 UTF-8 bytes. Names are scalar Unicode without C0/C1/DEL/BOM controls. A file picker or operating system may impose a shorter filename limit.
- Coordinates are ASCII decimals satisfying 0 ≤ start < end ≤ 9007199254740991. Signs, decimals, exponents and zero-width records are refused. Leading zero spelling stays in the matrix. Chromosome identifiers match case exactly; inputs may be unsorted.
- Intervals are 0-based, half-open [start,end). Touching endpoints do not overlap. Every duplicate or nested sample record counts separately; reference records are never deduplicated or sorted. The last columns follow the displayed sample order.
- Strict UTF-8 with one optional leading BOM and LF, CRLF or mixed line endings. Lone CR, invalid UTF-8, C0/C1/DEL controls and nonleading BOMs are refused. Blank/ASCII-whitespace lines and column-zero # comments are ignored for counts and retained in original sidecars.
- All downloads together allow 32 MiB: overlap-matrix.bed, full result-report.json and every original source sidecar. Exact size is checked before full counts, matrix text or report serialization is allocated. An over-budget task has no partial matrix, copy or downloads.
- The matrix uses tabs and LF, no header row, and retains every reference token plus one count per sample. The full report records every token/count, source names/SHA256/byte sizes/BOM/line endings/physical lines, sample record counts and complete export metadata. Originals retain all source bytes.
- Processing stays in a local Worker. The 10-second operation deadline covers worker loading, source reading, readiness and computation. Cancel and timeout terminate the worker; retry the same selected files and order. The screen previews up to 4,000 code points; copy and downloads are complete.
- This counts whole coordinate spans. It does not validate BED annotation columns, expand blocks, apply strand/fraction filters or infer genome assemblies. Appended sample columns may exceed 12; their meanings are counts rather than BED12 annotations. No file content or name is uploaded.