Build a quantitative DNA logo from an existing alignment
Choose the source format, gap rule and height model; verify complete column measurements and keep every SVG page, CSV, report and original.
Start from an alignment, not an image
Choose lines for one sequence per nonblank row, or FASTA for complete records whose headers begin with >. Select one original file or paste the alignment. If a file is selected, it takes priority even when text is still visible; clear the selection to use the text.
Sequences must already be equal length. ASCII spaces/tabs are removed and lowercase ACGT is normalized; this tool does not align sequences or resolve ambiguous bases. Keep the original so its source bytes can be checked against the report.
Set the height and gap model
For probability mode, each observed column has total height 1. For information mode, the fixed uniform ACGT background gives H = −Σ p log2 p, R = 2 − H and each base height p×R. There is no small-sample correction.
With pseudocount a, observed-column probabilities use (count+a)/(effectiveCount+4a). Raw counts remain unchanged. Ignore gaps accepts - and . without deleting columns; an all-gap column keeps null probabilities, entropy, information and heights even when a is positive.
| Choice | Meaning | Check |
|---|---|---|
| Probability | Share of observed bases after the chosen pseudocount | Each observed column totals 1 |
| Information | Probability weighted by information in bits | Column total equals R, up to 2 bits |
| Reject gaps | - or . refuses the whole alignment | Use ungapped input or deliberately change the rule |
| Ignore gaps | Gaps remain in place but are excluded from effectiveCount | Compare gapCount and effectiveCount per column |
Check the synthetic four-sequence example
Use the example shown below with information mode, reject gaps and pseudocount 0. Columns 1–3 contain only A, C and G respectively: each has entropy 0 and information 2 bits. Column 4 has one A, one G and two T.
For column 4 the probabilities are A=0.25, C=0, G=0.25 and T=0.5. Entropy is 1.5 bits and R is 0.5 bits, so heights are 0.125, 0, 0.125 and 0.25. These measurements come from this synthetic alignment; they do not reproduce an author attachment.
ACGT
ACGT
ACGA
ACGG
Save every column and page
Download the complete set after checking the sequence count, width and model in report.json. Column numbers are one-based. columns.csv contains four records per column, including zero-count bases and null-valued empty columns. All nonzero glyphs use the reported quantitative height.
SVG pages cover up to 500 columns each. A limited preview may show less than the complete result; download every page. Keep alignment.fasta for the normalized sequences and original.input for byte-exact source identity. Copy returns the entire compact report, exactly matching report.json.
| File | Purpose |
|---|---|
| report.json | Every sequence, model, source hash, column values and page count |
| columns.csv | All counts, effective/gap counts, probabilities, entropy, information and heights |
| alignment.fasta | Complete normalized sequence records |
| logo-NNN.svg | Every 500-column page with all nonzero glyphs |
| original.input | Exact effective source bytes |
- Check the sequence count and alignment width, then confirm the gap rule, pseudocount and chosen height model in report.json.
- Compare columns.csv counts and heights with the SVG; keep zero-count bases and all-gap null measurements visible in the data.
- Download every logo-NNN.svg page together with columns.csv, report.json, alignment.fasta and original.input; the preview does not include the whole delivery.
- After cancellation or timeout, retry the retained original File or text and the same options. Correct a refused alignment deliberately and retain that revised source separately.
Recover with a reproducible source
Input is limited to 10 MiB, 10,000 sequences, 3,000 columns and 9,999,999 cells; every limit applies. Complete files plus full report text are limited to 56 MiB. After a refusal, correct the source or explicitly reduce the analyzed alignment, then retain the resulting source and model.
The single 10-second deadline includes reading, loading, computation, result checks and the first display. Cancellation or timeout publishes no partial or late output. Retry the same File or text and options after stopping. Processing stays in this browser without persistent source storage. The original question’s textbook image has a reported error, so it is not a target to force the measurements to match.
References
- NKR — a logo from aligned DNA
The complete question and comments identify a textbook figure error. Calculate from the alignment; do not force the logo to match that image. No actual author attachment or completed result was obtained. This page’s four-sequence example is synthetic.
- Logomaker quantitative logo implementation
Supports reviewing count, probability and information transformations. This tool fixes a uniform DNA background and applies no small-sample correction; it does not claim identical defaults for another renderer.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
DNA sequence logoGenerate a quantitative ACGT logo from an existing alignment, with explicit probability or information heights, gap handling, full column measurements and exact source bytes.
Generate a quantitative ACGT logo from an existing alignment, with explicit probability or information heights, gap handling, full column measurements and exact source bytes.
Steps
- Choose lines or FASTA, then paste the existing aligned DNA or select one original file.
- Choose probability or information heights, a gap rule and a pseudocount from 0 through 100.
- Generate the logo and check column count, effective counts and complete column measurements.
- Keep every SVG page, column CSV, report, normalized FASTA and original; copy the full report for review.
Available options
- Input format
- One sequence per line · Multi-record FASTA
- Glyph height
- Probability (0–1) · Information (0–2 bits)
- Gap rule
- Reject - and . · Ignore - and .; keep columns
- Pseudocount per base
- 0
Finite 0–100, added separately to A/C/G/T; raw counts remain unchanged.
Capabilities and limits
- Paste an alignment or choose one UTF-8 file, up to 10 MiB. A selected file takes priority over pasted text; clear it to use the text. Filenames are limited to 512 UTF-8 bytes. No input content or name is uploaded.
- Choose one sequence per nonblank line or multi-record FASTA. ASCII spaces and tabs within sequences are removed and lowercase ACGT is converted to uppercase. FASTA fragments are concatenated; complete nonempty unique headers are retained, up to 256 UTF-16 units. One leading BOM, LF and CRLF are accepted; bare CR and malformed UTF-8 are refused.
- All sequences must already have the same length. Only ACGT is counted. With Ignore gaps, - and . remain at their original columns but are excluded from that column’s effective count. Reject gaps refuses them. IUPAC ambiguity, RNA, protein and alignment construction are outside this tool.
- At most 10,000 sequences, 3,000 columns and 9,999,999 sequence-by-column cells. All three limits apply together; 10,000 sequences of 1,000 columns exceed the cell limit.
- Pseudocount is a finite number from 0 through 100 added to each of A/C/G/T when computing probabilities. Raw counts stay unchanged. Information uses a fixed uniform background: R = 2 − H and height = probability × R, without small-sample correction. An all-gap column has null values and no glyph even with a positive pseudocount.
- Download report.json, columns.csv, alignment.fasta, every logo-NNN.svg page and original.input. Each SVG page covers at most 500 columns; no column or nonzero glyph is dropped. Column positions are one-based. The normalized FASTA and original bytes serve different purposes.
- Complete files plus full report text are limited to 56 MiB. Larger complete results are refused without shortening fields. Copy keeps the entire compact report; a limited screen preview does not replace the downloads.
- A single 10-second deadline covers source checks, file reading, loading, processing, result validation and the first result display. Cancel or timeout ends unfinished work without publishing a partial or late result. Keep the same source and options to retry.