Inspect local PDF links, comments and reading notes
Read actual PDF declarations and complete English Kindle records with explicit provenance, supported scope and full exports.
Choose the structure that answers the task
An unresponsive clickable PDF area needs its actual Link declaration. A reviewed report needs original annotation text and reply context. A device log needs complete reading records. These are three different inputs and deliveries; visible URLs, page text and line dedup cannot substitute for the underlying structures.
| Tool | Input task | Delivery |
|---|---|---|
| T294 | Clicking a PDF link does nothing | Page/rectangle and full target/action JSON/CSV |
| T295 | Convert reviewer comments into task records | Original text, author, raw dates and replies in JSON/CSV/Markdown |
| T296 | Back up local device reading notes | Exact-title grouping and source-to-retained records in JSON/CSV/Markdown |
Read a target without opening it
T294 scans real Link annotations. It resolves direct page targets, legacy Dests and the full Names/Dests tree using the original key bytes, declared Limits and unique page references. URI, GoTo, GoToR, Launch, Named and chained Next actions have explicit states. Other actions retain their type and declaration as unsupported. No action, URI or embedded script is executed, opened or fetched.
The checked three-page source contains 11 Links and two named destinations. Some resolve to pages 1–3, one has a missing target, a missing name remains dangling, and a JavaScript action stays an unsupported declaration. The downloaded CSV quotes full structured targets; a 2,000-character visible cell is only a preview.
- A written URI does not establish safety, network reachability or the reason a viewer ignored it.
- The same source can contain text URLs that are not clickable annotations; they do not increase Link count.
- Page rectangles are raw PDF points, not rotated or cropped screen coordinates.
- Malformed URI types, contradictory ownership, duplicate raw names, invalid NameTree ordering/Limits and cycles reject the whole inventory.
Keep review text and thread meaning
T295 exports the original plain Contents, author, original M/CreationDate strings, raw Rect/QuadPoints and IRT reply relationships. It does not reinterpret date timezones. Missing Contents is JSON null rather than invented text; RC presence is flagged and rich-only text is explicitly unsupported. The Markdown uses literal fenced blocks even when source text contains HTML or backticks.
The checked source has 17 supported records, 11 Links, one Widget and one Popup. It retains a multiline reply, a rich-only record and a dangling reply. Link, Widget and Popup have separately counted excluded roles. The complete source includes every supported subtype, including Redact; exporting a Redact annotation does not apply a redaction.
- Supported review types: Text, FreeText, Highlight, Underline, StrikeOut, Squiggly, Ink, Stamp, Caret, Line, Square, Circle, Polygon, PolyLine and Redact.
- Any other subtype, reused annotation identity, duplicate NM, conflicting P owner, malformed geometry or reply cycle fails the entire export.
- An absent or external reply object remains flagged as dangling; the tool does not fabricate its parent.
- This reads existing plain annotation text. It cannot reconstruct flattened review marks, recover highlighted sentences by OCR or execute rich HTML.
Organize whole clipping records instead of lines
T296 accepts an already copied UTF8 MyClippings.txt with English Highlight, Note and Bookmark metadata. A complete block contains the exact source book title, explicit positive page/location ranges, raw Added on text, body and a separate ========== line. That line is a reserved delimiter. BOM and CRLF/LF are reported; internal body line endings and Unicode remain unchanged.
The six-record BOM/CRLF example retains five records after exact dedup. An empty Bookmark is valid. The two excerpts at locations 40–45 and 42–47 both survive. Whole-record equality includes title, metadata, body and line endings; same location with different content, dates or whitespace does not overwrite anything. Original indexes map to retained IDs. Exact title strings define books without guessing an author split.
| Setting | Meaning |
|---|---|
| location | Within each exact book title, sort by explicit location start, otherwise page start; original index breaks ties |
| source | Keep source order within each exact book title |
| deduplicate=true | Keep one identical whole source block plus every original index |
| deduplicate=false | Retain every source record, including identical copies |
Download complete originals and inspect failures
The table is bounded to the first 200 records and 2,000 UTF16 characters per cell. Native copy contains a bounded processing summary. Downloads retain the full JSON records, all quoted CSV rows and literal Markdown bodies. CSV prefixes formula-like values with an apostrophe; JSON and Markdown preserve the originals. Dates remain raw source strings and no timezone is inferred.
The production checks use real 500-page sources with 10,000 Links and 10,000 review records, a separate 100,000-record clipping log, and an exact 10MiB single-note input. An independent pypdf/stdlib reader checks every source record, reference, target, reply, CSV cell and Markdown body. Real core progress cancellation followed by the same input recovers the full delivery; browser interaction acceptance is separately maintained by the controller.
- Each PDF: 30MiB, 500 pages, 200,000 parsed objects, depth 64, 10,000 annotations, 16MiB extracted strings and 4,096 bytes per PDF name. T294 adds 10,000 destinations and 1,000 NameTree nodes. JSON/CSV for links or JSON/CSV/Markdown for reviews total at most 20MiB. Object/XRef streams total 80MiB decoded; plain/Flate/LZW/ASCII85/ASCIIHex/RunLength, up to four filters, Predictor=1 only (LZW EarlyChange=0/1). Both cases of legal #hex names are parsed without dropping records.
- Clippings: one 10MiB UTF8 log and 100,000 complete records; three complete outputs total at most 40MiB. The source and output limits both apply; a short but highly repetitive metadata structure may still exceed the output budget.
- Empty logs and PDFs without matching records yield valid empty lists. Unsupported locale/type, bad UTF8, reversed/repeated ranges and missing delimiters reject clippings with record position.
- Failure or cancellation produces no partial success report. Correct the indicated source structure and retry; originals stay unchanged.
- No device/account access, URL requests, cloud sync, Kindle reimport, missing-log recovery, source-link repair or comment removal is performed.
References
- User task: inspect the original target of an unresponsive PDF hyperlink
Original question body read 2026-10-07; a written target is not proof of network reachability.
- User task: export multiline PDF review comments into a task system
Original body read 2026-10-07; separate real comments without losing internal line breaks.
- User task: back up local notes, highlights and bookmarks
Original body and relevant local MyClippings answer read 2026-10-07; account sync and reimport are outside this workflow.
Tools in this category
Expand a tool to see its steps, options and supported formats, then open its workspace.
PDF clickable link target inventoryRead real PDF Link annotations, declared actions and named destinations into complete local JSON and CSV.
Inspect actual clickable annotations, including missing, dangling and explicitly unsupported targets. URI, JavaScript and launch declarations remain inert text; nothing is opened, executed, requested or repaired. Rectangles use raw PDF points, not rotated screen coordinates.
Steps
- Choose a local PDF.
- Review target states and raw page rectangles.
- Download complete JSON/CSV before investigating a declared target elsewhere.
Capabilities and limits
- One unsigned, unencrypted PDF up to 30MiB and 500 pages; 200,000 parsed objects, depth 64, 10,000 annotations and destinations, 1,000 NameTree nodes, 16MiB extracted strings and 4,096 bytes per PDF name. Full JSON/CSV combined 20MiB. Object/XRef streams total 80MiB decoded; plain/Flate/LZW/ASCII85/ASCIIHex/RunLength, up to four filters, Predictor=1 only (LZW EarlyChange=0/1). Both cases of legal #hex names are parsed without dropping records.
- Resolve direct targets, legacy Dests and full Names/Dests trees using original byte ordering and limits. First release declares URI, GoTo, GoToR, Launch, Named and Next; other actions retain original type/declaration as unsupported. Cycles, conflicting page ownership, duplicate names, malformed trees and invalid URI types fail the whole inventory.
- Preview only the first 200 rows and 2,000 UTF16 characters per cell. Complete downloads retain targets. A written URI is not evidence of network reachability or safety; text URLs outside Link annotations are not counted.
PDF review comment and reply exportExport original PDF review text, authors, raw dates, geometry and reply relationships as complete JSON, CSV and Markdown.
Read actual review annotations without changing the PDF. Keep multiline plain Contents, raw author/date declarations and reply context. Missing Contents remains null; rich markup presence is flagged, never executed or guessed into text. Highlight artwork and flattened page text are not OCR inputs.
Steps
- Choose the reviewed local PDF.
- Review plain text, missing-content flags and reply status.
- Download complete JSON, quoted CSV and literal Markdown for your task system.
Capabilities and limits
- One unsigned, unencrypted PDF up to 30MiB and 500 pages; 200,000 objects, depth 64, 10,000 annotations, extracted strings 16MiB and PDF names 4,096 bytes each. JSON/CSV/Markdown combined 20MiB. Object/XRef streams total 80MiB decoded; plain/Flate/LZW/ASCII85/ASCIIHex/RunLength, up to four filters, Predictor=1 only (LZW EarlyChange=0/1). Both cases of legal #hex names are parsed without dropping records.
- Supported: Text, FreeText, Highlight, Underline, StrikeOut, Squiggly, Ink, Stamp, Caret, Line, Square, Circle, Polygon, PolyLine and Redact. Link, Widget and Popup are excluded with exact counts. Any other subtype fails the whole export.
- Raw PDF Rect/QuadPoints are source coordinates, not visible screen coordinates. Duplicate identities/names, contradictory owner pages, malformed geometry and reply cycles fail. Dangling replies stay flagged; 200 preview rows and 2,000 UTF16 characters per cell do not truncate downloads.
Kindle local clipping organizerGroup complete English MyClippings records by original book title and explicit location, keeping multiline notes and exact duplicate provenance.
Choose a UTF8 MyClippings.txt already copied from your device. Parse Highlight, Note and Bookmark metadata; keep page/location ranges and the original Added on text without inferring a timezone. Group exact source titles and optionally remove identical whole records; overlapping or different excerpts remain separate.
Steps
- Copy your own MyClippings.txt from the device and select it.
- Choose explicit range sorting or source order and optional exact dedup.
- Download complete grouped JSON, CSV and literal Markdown; keep the original log.
Available options
- Ordering
- Explicit location/page range · Original order within each book
- Remove exact whole-record duplicates
- On by default
Capabilities and limits
- One UTF8 file up to 10MiB and 100,000 complete records; complete JSON/CSV/Markdown combined 40MiB. English metadata only; unsupported locale/type, malformed or reversed ranges, repeated declarations, bad UTF8 and missing delimiters fail with source position.
- The exact line ========== is a reserved record delimiter. Original BOM and CRLF/LF are reported; internal body line endings and Unicode remain unchanged. Exact dedup includes raw metadata and line endings, retaining every original-to-kept mapping. Title/author ambiguity is not guessed.
- Preview only the first 200 records and 2,000 UTF16 characters per cell; download all records. CSV guards formula-like values, JSON/Markdown keep originals. No device/account access, cloud sync, Kindle reimport or missing-log recovery. Empty logs and empty Bookmark bodies are valid.