RDF dataset canonical comparison
Compare two local RDF N-Quads datasets using RDFC-1.0 canonicalization, including renamed blank nodes and named graphs.
- 1Add input
- 2Adjust settings
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Find out whether two RDF exports describe the same dataset despite different blank-node labels. Download both canonical datasets and inspect additions and removals from their separately canonicalized line sets.
How to use this tool
- Paste strict N-Quads on both sides, or open the left local export and paste the right one.
- Run canonical comparison and check equivalence, unique quad counts and additions/removals.
- Download left and right canonical N-Quads plus the comparison report. Treat blank-node-heavy differences as a dataset comparison, not a minimal edit plan.
Supported inputs and limits
One UTF-8 file or pasted left dataset plus a pasted right dataset, combined up to 2 MiB and 10,000 parsed quads. Empty datasets are valid. Repeated identical quads have set semantics and duplicate counts are reported. Selected left file takes precedence.
Strict RDF 1.1 N-Quads only: absolute IRIs, blank nodes, datatype/language literals and named graphs. Turtle prefixes, RDF/XML, RDF-star, relative IRIs and invalid quad roles reject. Language tags are normalized to lowercase by the parser; lexical IRI and datatype meanings are retained. No inference or automatic IRI alignment.
Canonicalization is explicitly RDFC-1.0 with SHA-256, not an undeclared algorithm alias. Each connected blank-node component is limited to 256 nodes, deep iterations to 10,000, and work to a 10-second deadline. Exceeding a work budget rejects atomically. Cancellation uses the dedicated worker and algorithm abort signal. Canonical downloads are combined up to 10 MiB.
Added and removed lines compare separately canonicalized datasets, not minimum edit distance or persistent blank-node identity. A small graph edit may relabel many blank nodes and enlarge the displayed difference. This tool does not certify RDF signatures. Preview shows the first 200 difference rows; JSON and canonical N-Quads downloads are complete.
Worked example
Example input
_:alice <https://example.com/name> "Alice"@en <https://example.com/team> .
Example options
{"secondary": "_:person <https://example.com/name> \"Alice\"@en <https://example.com/team> .\n", "params": {}}Example output
{
"algorithm": "RDFC-1.0",
"equivalent": true,
"leftQuads": 1,
"rightQuads": 1,
"leftDuplicates": 0,
"rightDuplicates": 0,
"removed": [],
"added": [],
"comparison": "separately-canonicalized-set-difference"
}When something does not work
Fix N-Quads syntax or export unsupported RDF formats first. If a symmetry/work budget is exceeded, simplify or partition only when the dataset semantics permit it. Rerun the example after correction; failed runs provide no partial canonical output.
Frequently asked questions
Do renamed blank nodes count as a change?
No, when the datasets are isomorphic and complete within the work budgets, RDFC-1.0 assigns a common canonical form.
Why can a small edit show many changed lines?
The datasets are canonicalized separately. Changed graph structure can change canonical blank-node labels.
Can it compare Turtle or RDF/XML?
Export those formats as strict RDF 1.1 N-Quads first. This task deliberately accepts one interchange syntax.