Pairwise Spearman needs ranks from the same valid rows
Use a finite example with ties and missing cells to explain why pair filtering changes ranks, and how repeated labels and complete results remain auditable.
The label request came from a resolved question
The author wanted matrix labels without a dataframe name and selected Age_55_64 twice. The saved discussion shows the problem was resolved in R and contains no complete original numeric input. This tool retains labels and selection identity while adding Spearman, effective N, status and SVG; it cannot claim to reproduce the author’s numbers.
Removing rows changes tie-rank positions
The fixture’s first column is 1,2,2,4,NA,6. Paired with GDP, one row containing 2 is excluded because GDP is missing, leaving 1,2,4,6 with ranks 1,2,3,4. Ranking the whole column first gives the two values of 2 rank 2.5; deleting a row afterward does not repair that set difference.
| Step | Behavior |
|---|---|
| 1 | Filter valid rows using the declared deletion policy. |
| 2 | Compute each variable’s average tie ranks within that subset. |
| 3 | Compute Pearson correlation of the ranks and retain N and status. |
A matching name need not mean the same physical column
Raw labels can match, and repeated choices can reference one physical column. JSON retains the selection ordinal, physical index and raw header so the recipient can identify the source without relying on temporary renaming.
- Repeated choices still produce complete rows and columns.
- Every cell’s effective N and status travel with its coefficient.
Zero correlation and unavailable calculation have different meanings
A constant column, fewer than two valid rows or detected numerical underflow can leave no coefficient. A null status explains why; zero is a computed coefficient. Scaling and compensated sums reduce overflow risk, but finite binary numbers still have precision boundaries, so the original strings remain complete.
Save complete evidence after the finite preview
The first 10×10 cells support inspection; full CSV, JSON, the original file and fully labelled SVG provide the complete handoff. Row count, column choices and output size share the same limits. If a task exceeds a joint budget or the 30-second deadline, no partial matrix is returned; reduce the task while retaining the original file and settings.
References
- Original question
The seven saved complete public posts were read. The author resolved the label issue in R. No complete original numeric CSV is available; browser preference and market size are unknown.
- jStat 1.9.6 core and MIT license
Average tie ranks and correlation semantics are retained with scoped changes for bounds, actual comparison counting and stable arithmetic. The complete MIT license is included.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
Column correlation matrixCompute Pearson or Spearman correlations from a complete CSV, preserving original labels, repeated selections, each pair’s effective N and status.
Select economic, social or other numeric variables by their physical column numbers. Preserve raw labels and repeated choices, compute every variable pair, and use N and status to explain unavailable cells. Pairwise and listwise deletion can produce different valid-row sets.
Steps
- Choose the complete UTF-8 file and check its delimiter and header setting.
- Enter physical column numbers, choose Pearson or Spearman, and declare deletion policy and exact missing markers.
- Inspect coefficients, N and statuses in the finite preview, then download the full matrix, effective N, status JSON, SVG, original, settings and license.
Available options
- Numeric column numbers
- 1,2
Comma-separated physical column numbers starting at 1, such as 1,2,3,1. Repeated choices retain distinct selection ordinals.
- Correlation method
- Pearson · Spearman
Spearman computes average tie ranks separately within each pair-valid subset.
- Missing-data policy
- Pairwise deletion · Listwise deletion
Pairwise uses rows valid for that pair; listwise uses rows valid for every selected physical column.
- Input delimiter
- Comma · Semicolon · Tab · Pipe
- Header row
- First row is a header · No header
Duplicate or empty headers remain unchanged; no-header input receives labels such as column_1.
- Missing-value markers
- Enter as needed
One per line, case-sensitive after trimming. Empty cells are always missing.
Capabilities and limits
- One nonempty UTF-8 CSV/TSV/TXT file: at most 20 MiB, 50,000 data rows and 2,000,000 cells including a header when present. Rows must have equal width; UTF-8 BOM, quotes and quoted line breaks are supported.
- Select 1–64 ordinals, including repeats. Selected nonmissing cells must be finite binary64 numbers in decimal or scientific notation. Other text, infinity and source underflow refuse the whole task. Original numeric strings and unselected columns remain complete.
- Up to 20 missing markers, 100 characters each and 2,048 characters overall. Fewer than two valid rows or constant columns produce a null coefficient and explicit status, as do detected arithmetic underflow, lost distinctions or nonfinite intermediates. Only an excess beyond ±1 within 32 Number.EPSILON is corrected and recorded.
- Joint budgets: at most 100,000,000 pair-row work units and 20,000,000 actual sort comparisons; files plus full text 32 MiB, compact typed JSON 16 MiB, their aggregate 48 MiB, binary plus metadata transport 64 MiB and conservative estimated memory 512 MiB. Axis ceilings do not promise simultaneous maximum capacity.
- A 30-second absolute deadline covers native file validation, reading, Worker loading, calculation, full-result validation, cleanup and first UI publication after await. Failure or cancellation publishes no partial result. The view shows at most 10×10 cells and 120 characters per label; downloads retain every cell and full label.