Extract PDF attachments
Recover embedded files from PDF attachment structures and download their original decoded payloads with a filename map.
- 1Add input
- 2Review and run
- 3Get your result
Tool input and files are processed in this browser without being uploaded.
Before you start
Recover invoice XML, documents or other files stored inside a PDF. The attachment list records original and safe download names, sizes and locations; each file and a ZIP are available locally.
How to use this tool
- Upload the PDF and extract its standard embedded file attachments.
- Review the original-to-download name mapping, decoded sizes and attachment locations. External references are listed without downloading their targets.
- Download individual files or the ZIP, then open the expected format in your own application. A PDF without supported attachment entries returns an empty report.
Supported inputs and limits
One unencrypted PDF up to 30 MiB and 300 pages; up to 100 file specifications, 10 MiB per decoded attachment and 30 MiB total. Supports standard EmbeddedFiles NameTrees and page FileAttachment annotations with embedded Filespec EF streams.
Only unfiltered streams or a single FlateDecode filter without DecodeParms are supported. External file references are reported and never fetched. This extracts file attachments rather than page images; files are not opened or executed. Unsafe or duplicate filenames are renamed with an explicit mapping.
Worked example
Example input
A PDF with a 52-byte UTF-8 invoice XML, an empty text attachment, a 5-byte binary attachment and one external file reference.
Example options
{}Example output
Three files totaling 57 bytes, one ZIP and a JSON filename/location map; the external file is not fetched.
When something does not work
Use an unencrypted PDF containing standard embedded attachments. Unsupported encodings, predictors, malformed attachment trees or exceeded decoded limits are rejected. If the report is empty, inspect whether the PDF actually embeds the files or merely links to them.
Frequently asked questions
Can this recover the XML stored in a PDF invoice?
Yes, when the XML is a standard embedded file using a supported stream encoding. The extracted bytes are the decoded payload stored in the PDF.
Why are download names different?
Directory paths, unsafe characters, reserved filenames and collisions are normalized for local downloads. The JSON report preserves the original name and gives the exact output name.
Will it download an external attachment or run a script?
No. External references remain report entries only. Embedded files are offered as downloads without being opened or executed; page images are outside this tool's task.