Choose probability or information heights for a DNA logo
Compare frequencies and bits, effective counts and empty columns so the figure follows the actual alignment and explicit model.
A taller stack can answer a different question
Probability shows composition: after a chosen pseudocount, how much of a column is A, C, G or T? Information height also weights that composition by departure from a fixed uniform background. Choosing the mode changes the vertical quantity, so keep it with the figure.
In the synthetic ACGT/ACGT/ACGA/ACGG alignment, column 4 has probabilities 0.25, 0, 0.25, 0.5. Probability mode totals 1; information mode totals 0.5 bits and gives heights 0.125, 0, 0.125, 0.25. The relative shares match, while the stack’s total height changes.
| Composition | Probability total | Information total |
|---|---|---|
| One base in every sequence | 1 | 2 bits |
| Two bases equally frequent | 1 | 1 bit |
| All four bases equally frequent | 1 | 0 bits |
A zero-height column and a missing column differ
An observed uniform ACGT column has entropy 2 bits and information 0. Its information-mode glyph heights are numeric zero. An all-gap column has effectiveCount 0, null measurements and no glyph; it remains at its original column position.
Ignore gaps changes the denominator per column, not the alignment width. Compare effectiveCount with gapCount before reading a tall stack as strong evidence. A column with only one observed base is still a small observation; this tool does not apply a small-sample correction.
A pseudocount is a visible modeling choice
Pseudocount a is added separately to A/C/G/T for probability calculation. It leaves source counts unchanged but spreads probability toward a uniform distribution. Use 0 for unadjusted observed frequencies or enter a finite value through 100 deliberately.
The model uses a uniform background of 0.25 for each base. A nonuniform genomic background, ambiguity decoding, RNA, proteins and automatic alignment are outside this task. Keep the complete model and normalized sequences with the measurements instead of transferring only the figure.
Check a figure against its data
The original learner asked why a logo did not match a textbook example. The complete comments identify an error in that image; the feedback supports checking the actual columns instead of adjusting the calculation to resemble it. No actual author attachment or completed result was obtained.
Use a practical handoff: choose the format and model, run the full alignment, inspect several known columns, and download columns.csv alongside report.json and every SVG page. In the synthetic example, column 4 must retain one A, one G and two T before any pseudocount.
- Record the mode, gap rule and pseudocount before comparing figures.
- Check raw counts, effectiveCount and gapCount before checking glyph heights.
- Keep all pages; the screen preview is a separate bounded view.
- Retain both normalized FASTA and byte-exact original with the source hash.
Keep limits and recovery visible
The input ceilings are 10 MiB, 10,000 sequences, 3,000 columns and 9,999,999 cells; outputs remain complete within the 56 MiB files-plus-report-text budget. These are protective limits, not measured browser-heap promises.
A cancellation, timeout or output refusal yields no partial figure. Keep the source and parameters for a fresh retry; if you reduce the analyzed alignment, record that change. The logo describes an explicit mathematical model of the input, without establishing biological function.
References
- NKR — a logo from aligned DNA
The complete question and comments identify a textbook figure error. Calculate from the alignment; do not force the logo to match that image. No actual author attachment or completed result was obtained. This page’s four-sequence example is synthetic.
- Logomaker quantitative logo implementation
Supports reviewing count, probability and information transformations. This tool fixes a uniform DNA background and applies no small-sample correction; it does not claim identical defaults for another renderer.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
DNA sequence logoGenerate a quantitative ACGT logo from an existing alignment, with explicit probability or information heights, gap handling, full column measurements and exact source bytes.
Generate a quantitative ACGT logo from an existing alignment, with explicit probability or information heights, gap handling, full column measurements and exact source bytes.
Steps
- Choose lines or FASTA, then paste the existing aligned DNA or select one original file.
- Choose probability or information heights, a gap rule and a pseudocount from 0 through 100.
- Generate the logo and check column count, effective counts and complete column measurements.
- Keep every SVG page, column CSV, report, normalized FASTA and original; copy the full report for review.
Available options
- Input format
- One sequence per line · Multi-record FASTA
- Glyph height
- Probability (0–1) · Information (0–2 bits)
- Gap rule
- Reject - and . · Ignore - and .; keep columns
- Pseudocount per base
- 0
Finite 0–100, added separately to A/C/G/T; raw counts remain unchanged.
Capabilities and limits
- Paste an alignment or choose one UTF-8 file, up to 10 MiB. A selected file takes priority over pasted text; clear it to use the text. Filenames are limited to 512 UTF-8 bytes. No input content or name is uploaded.
- Choose one sequence per nonblank line or multi-record FASTA. ASCII spaces and tabs within sequences are removed and lowercase ACGT is converted to uppercase. FASTA fragments are concatenated; complete nonempty unique headers are retained, up to 256 UTF-16 units. One leading BOM, LF and CRLF are accepted; bare CR and malformed UTF-8 are refused.
- All sequences must already have the same length. Only ACGT is counted. With Ignore gaps, - and . remain at their original columns but are excluded from that column’s effective count. Reject gaps refuses them. IUPAC ambiguity, RNA, protein and alignment construction are outside this tool.
- At most 10,000 sequences, 3,000 columns and 9,999,999 sequence-by-column cells. All three limits apply together; 10,000 sequences of 1,000 columns exceed the cell limit.
- Pseudocount is a finite number from 0 through 100 added to each of A/C/G/T when computing probabilities. Raw counts stay unchanged. Information uses a fixed uniform background: R = 2 − H and height = probability × R, without small-sample correction. An all-gap column has null values and no glyph even with a positive pseudocount.
- Download report.json, columns.csv, alignment.fasta, every logo-NNN.svg page and original.input. Each SVG page covers at most 500 columns; no column or nonzero glyph is dropped. Column positions are one-based. The normalized FASTA and original bytes serve different purposes.
- Complete files plus full report text are limited to 56 MiB. Larger complete results are refused without shortening fields. Copy keeps the entire compact report; a limited screen preview does not replace the downloads.
- A single 10-second deadline covers source checks, file reading, loading, processing, result validation and the first result display. Cancel or timeout ends unfinished work without publishing a partial or late result. Keep the same source and options to retry.