An overlap count depends on records and sample order
Understand adjacency, repeated intervals, opaque fields and why complete originals matter.
Record counts answer a specific question
Two nested records can cover the same bases and still contribute two overlaps. A union-coverage calculation would merge their geometry and answer a different question. This matrix retains each record’s contribution.
Repeated reference rows also remain repeated. Their counts may be identical, but their opaque annotations or original positions can matter to the analysis that consumes the result.
chr1 0 1000
chr1 1000 2000
A count column needs provenance
Changing sample order changes the meaning of matrix columns without changing the reference coordinates. A no-header matrix needs an accompanying list of ordered sources. The report retains sample roles, complete names and SHA256 values to make that order reviewable.
Normalizing spaces to tabs makes the matrix easier to read; it does not preserve the source formatting. The original sidecars carry the full comments, separators, BOM and line endings. Use those bytes to investigate a disputed result.
| Observation | Result in this example | Meaning |
|---|---|---|
| sample-001 overlap records | Reference [0,1000): 2 | [20,50) and [30,40) each add 1; this is the matrix field. |
| sample-001 union-covered length | Reference [0,1000): 30 bases | [30,40) is contained in [20,50); the matrix does not output coverage length. |
| sample-002 overlap records | Across the reference rows: 2, then 0 | The second count column comes from the two records below, not from filename sorting. |
| Sample column order | Second row: 1,0; swapped: 0,1 | Reference fields stay unchanged; keep the ordered sources in the report with the matrix. |
chr1 20 50
chr1 30 40
chr1 1500 1900
Span overlap is a finite analysis profile
BED extra columns can describe blocks, strand and annotations, but these tokens are opaque in this tool. Counting the first-three-column span is appropriate only when that is the question you intend to answer. A block-aware or strand-filtered task needs its own analysis workflow.
A complete result at the joint download limit is useful; a truncated matrix is misleading. The tool measures the complete matrix, report and all originals before allocating the full result. If the work exceeds a limit, correct or split the inputs and retain the boundary between runs.
chr1 68 786
chr1 899 987
- Before running, confirm that the question is whole-span overlap record counts. Use an appropriate workflow for covered length, block expansion or strand filters.
- For unexpected zeros, check genome assembly and chromosome case, then half-open endpoints. For high counts, check whether nested and repeated records should each contribute.
- When sharing results, retain the matrix, report and every original sidecar. Trace each count column using the report’s order, names and SHA256 values.
- If limits require splitting by sample, use the same complete reference in every run. Keep each report and sample order, then align columns by reference row; do not silently truncate output.
Keep the complete result with its interpretation
Running reference.bed with sample-001.bed, then sample-002.bed produces the complete matrix below. The first-row counts 2,2 do not imply equal covered length; the second-row counts 1,0 identify different record counts for that span.
chr1 0 1000 2 2
chr1 1000 2000 1 0
References
- Reference-by-sample overlap matrix request
Complete source question and both answers retained in the admission record. No author success, browser preference or demand-volume claim.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
BED overlap count matrixCount each sample interval that overlaps each reference span, including duplicate and nested records. Keep every reference field, the chosen sample order, exact originals and a complete source report.
Count each sample interval that overlaps each reference span, including duplicate and nested records. Keep every reference field, the chosen sample order, exact originals and a complete source report.
Steps
- Choose the reference file first, then the sample files.
- Move samples up/down until their order matches the intended matrix columns.
- Run and inspect reference/sample totals and complete download sizes.
- Copy the complete matrix or download the matrix, report and every original. Verify coordinate assembly compatibility in your analysis workflow.
Capabilities and limits
- Choose one reference BED file and 1–20 sample BED files in explicit column order. All original sources together allow 4 MiB and 100,000 interval records. A physically empty file is refused; a nonempty blank/comment-only sample contributes zeros. The reference must contain at least one interval.
- Each record has 3–12 nonempty tokens, separated by ASCII spaces or tabs. The first three are chromosome, start and end; every remaining token is opaque and retained. Leading/trailing ASCII separators normalize in the matrix. Quotes and backslashes are literal; empty fields or whitespace inside a field are unsupported.
- Chromosome and source labels allow 512 UTF-8 bytes; each coordinate or extra token allows 65,536 UTF-8 bytes. Names are scalar Unicode without C0/C1/DEL/BOM controls. A file picker or operating system may impose a shorter filename limit.
- Coordinates are ASCII decimals satisfying 0 ≤ start < end ≤ 9007199254740991. Signs, decimals, exponents and zero-width records are refused. Leading zero spelling stays in the matrix. Chromosome identifiers match case exactly; inputs may be unsorted.
- Intervals are 0-based, half-open [start,end). Touching endpoints do not overlap. Every duplicate or nested sample record counts separately; reference records are never deduplicated or sorted. The last columns follow the displayed sample order.
- Strict UTF-8 with one optional leading BOM and LF, CRLF or mixed line endings. Lone CR, invalid UTF-8, C0/C1/DEL controls and nonleading BOMs are refused. Blank/ASCII-whitespace lines and column-zero # comments are ignored for counts and retained in original sidecars.
- All downloads together allow 32 MiB: overlap-matrix.bed, full result-report.json and every original source sidecar. Exact size is checked before full counts, matrix text or report serialization is allocated. An over-budget task has no partial matrix, copy or downloads.
- The matrix uses tabs and LF, no header row, and retains every reference token plus one count per sample. The full report records every token/count, source names/SHA256/byte sizes/BOM/line endings/physical lines, sample record counts and complete export metadata. Originals retain all source bytes.
- Processing stays in a local Worker. The 10-second operation deadline covers worker loading, source reading, readiness and computation. Cancel and timeout terminate the worker; retry the same selected files and order. The screen previews up to 4,000 code points; copy and downloads are complete.
- This counts whole coordinate spans. It does not validate BED annotation columns, expand blocks, apply strand/fraction filters or infer genome assemblies. Appended sample columns may exceed 12; their meanings are counts rather than BED12 annotations. No file content or name is uploaded.