# All Pages
Source: https://developer.pdf.co/all
Complete index of every PDF.co documentation page, grouped by section.
A complete directory of every PDF.co documentation page, grouped by section — a crawlable map of the docs. Jump to any topic below.
## API Reference
* [Getting Started](/api)
* [Get Account Balance Info](/api/account-balance-info)
* [AI Invoice Parser](/api/ai-invoice-parser)
* [Understanding Sync and Async Modes](/api/async-and-sync-mode)
* [Barcodes Generator](/api/barcode/generate)
* [Barcodes Overview](/api/barcode/overview)
* [Barcodes Reader](/api/barcode/read)
* [Excel to CSV](/api/convert-from-excel/csv)
* [Excel to HTML](/api/convert-from-excel/html)
* [Excel to JSON](/api/convert-from-excel/json)
* [Excel to PDF](/api/convert-from-excel/pdf)
* [Excel to Text](/api/convert-from-excel/text)
* [Excel to XML](/api/convert-from-excel/xml)
* [Document Classifier](/api/document-classifier)
* [Document Parser Overview](/api/documentparser/overview)
* [Parse Document](/api/documentparser/parser)
* [List All Templates](/api/documentparser/templates)
* [Retrieve Template by ID](/api/documentparser/templates-id)
* [Extract Data from Email File](/api/email/decode)
* [Extract Email Attachment](/api/email/extract-attachments)
* [Send Email with File](/api/email/send)
* [File Download](/api/file-download)
* [Delete Temporary File](/api/file-upload/delete)
* [Generate Pre-signed URL](/api/file-upload/generate-presigned-url)
* [Get MD5 Hash of File by URL](/api/file-upload/hash)
* [File Upload Overview](/api/file-upload/overview)
* [Upload Small File](/api/file-upload/upload)
* [Upload File Using Base64](/api/file-upload/upload-base64)
* [Upload File via Pre-signed URL](/api/file-upload/upload-presigned-url-put)
* [Upload File from URL \[GET\]](/api/file-upload/upload-url-get)
* [Upload File from URL \[POST\]](/api/file-upload/upload-url-post)
* [PDF Forms Info Reader](/api/forms/info-reader)
* [Background & Job Check](/api/job-check)
* [Language Support](/api/language-support)
* [Merge PDF](/api/merge/pdf)
* [Merge Various Document Type](/api/merge/various-files)
* [PDF Add](/api/pdf-add)
* [Make Text Searchable](/api/pdf-change-text-searchable/searchable)
* [Make Text Unsearchable](/api/pdf-change-text-searchable/unsearchable)
* [PDF Compress](/api/pdf-compress)
* [PDF Delete Pages](/api/pdf-delete-pages)
* [Extract Attachment](/api/pdf-extract-attachments)
* [PDF Find Text](/api/pdf-find/basic)
* [Find Text in Table with AI](/api/pdf-find/table)
* [PDF from CSV](/api/pdf-from-document/csv)
* [PDF from DOC](/api/pdf-from-document/doc)
* [PDF from Email](/api/pdf-from-email)
* [PDF from HTML](/api/pdf-from-html/convert)
* [PDF from HTML Template](/api/pdf-from-html/convert-from-template)
* [Return HTML Template by ID](/api/pdf-from-html/template-id)
* [Return All Templates](/api/pdf-from-html/templates)
* [PDF from Image](/api/pdf-from-image)
* [PDF from URL](/api/pdf-from-url)
* [PDF Info Reader](/api/pdf-info-reader)
* [Add Password to PDF](/api/pdf-password/add)
* [Remove Password from PDF](/api/pdf-password/remove)
* [Auto-rotate Pages with AI](/api/pdf-rotate/auto)
* [Rotate Selected Pages](/api/pdf-rotate/basic)
* [PDF Search and Delete Text](/api/pdf-search-text-and-delete)
* [Search and Replace with Image](/api/pdf-search-text-and-replace/image)
* [Search and Replace with Text](/api/pdf-search-text-and-replace/text)
* [Split PDF](/api/pdf-split/by-pages)
* [Split PDF by Text Search](/api/pdf-split/by-text-search-or-barcode)
* [PDF to CSV](/api/pdf-to-csv)
* [PDF to XLS](/api/pdf-to-excel/xls)
* [PDF to XLSX](/api/pdf-to-excel/xlsx)
* [PDF to HTML](/api/pdf-to-html)
* [PDF to JPG](/api/pdf-to-image/jpg)
* [PDF to PNG](/api/pdf-to-image/png)
* [PDF to TIFF](/api/pdf-to-image/tiff)
* [PDF to WEBP](/api/pdf-to-image/webp)
* [PDF to JSON](/api/pdf-to-json/basic)
* [PDF to JSON with AI](/api/pdf-to-json/with-ai)
* [PDF to Text](/api/pdf-to-text/basic)
* [PDF to Text (Simple)](/api/pdf-to-text/simple)
* [PDF to XML](/api/pdf-to-xml)
* [Postman](/api/postman)
* [Profiles](/api/profiles)
* [Response Codes](/api/response-codes)
* [Credits per API Function](/api/credits-per-api-function)
* [URL Input and Request Limits](/api/url-input-and-request-limits)
* [Webhook and Callbacks](/api/webhooks)
## Integrations
* [PDF.co Integrations](/integrations)
* [Airtable](/integrations/airtable)
* [Google Apps Script](/integrations/google-apps-script)
* [Add Password and Security into PDF](/integrations/make/add-security-to-pdf)
* [Add Text and Images To a PDF](/integrations/make/add-text-images-formfields-to-pdf)
* [AI Invoice Parser](/integrations/make/ai-invoice-parser)
* [Generate a Barcode](/integrations/make/barcode-generate)
* [Read a Barcode](/integrations/make/barcode-read)
* [Compress and Optimize PDF](/integrations/make/compress-pdf)
* [Convert from PDF](/integrations/make/convert-from-pdf)
* [Convert into PDF](/integrations/make/convert-to-pdf)
* [Create Fillable PDF Form](/integrations/make/create-fillable-pdf-form)
* [Document Classifier](/integrations/make/document-classifier)
* [Parse a Document](/integrations/make/document-parser)
* [Fill a PDF Form](/integrations/make/fill-pdf-form)
* [Getting Started with Make](/integrations/make/getting-started)
* [Convert HTML to PDF](/integrations/make/html-to-pdf)
* [Convert from Images into PDF](/integrations/make/images-to-pdf)
* [Integrating File Sources with PDF.co](/integrations/make/input-file-sources)
* [Job Check](/integrations/make/job-check)
* [Make Webhooks Integration with PDF.co](/integrations/make/make-webhooks)
* [Merge a PDF](/integrations/make/merge)
* [Get PDF Information](/integrations/make/pdf-info)
* [Convert from PDF into Images](/integrations/make/pdf-to-images)
* [Make PDF.co API Call](/integrations/make/pdfco-api-call)
* [Remove Password and Security from PDF](/integrations/make/remove-security-from-pdf)
* [Search and Delete Found Text in PDF](/integrations/make/search-and-delete-text)
* [Search and Replace Text in PDF](/integrations/make/search-and-replace-text)
* [Search and Replace With Image in PDF](/integrations/make/search-and-replace-with-image)
* [Search Text in PDF](/integrations/make/search-text)
* [Send Email With Attachments](/integrations/make/send-email-with-attachments)
* [Split a PDF](/integrations/make/split-pdf)
* [Upload a File](/integrations/make/upload-file)
* [Getting Started with Microsoft Power Automate](/integrations/microsoft-power-automate/getting-started)
* [Integrating File Sources with PDF.co](/integrations/microsoft-power-automate/input-file-sources)
* [Add Text or Images to PDF](/integrations/n8n/add-text-image-to-pdf)
* [AI Invoice Parser](/integrations/n8n/ai-invoice-parser)
* [Barcode Generator](/integrations/n8n/barcode-generator)
* [Barcode Reader](/integrations/n8n/barcode-reader)
* [Compress PDF](/integrations/n8n/compress-pdf)
* [Convert PDF to Anything](/integrations/n8n/convert-from-pdf)
* [Convert Anything to PDF](/integrations/n8n/convert-to-pdf)
* [PDF.co API and n8n Integration Guide](/integrations/n8n/custom-api-call)
* [Delete PDF Pages](/integrations/n8n/delete-page-in-pdf)
* [Fill a PDF Form](/integrations/n8n/fill-a-pdf-form)
* [Get PDF Information & Form Fields](/integrations/n8n/get-pdf-information)
* [Getting Started with n8n](/integrations/n8n/getting-started)
* [Make PDF Searchable/Unsearchable](/integrations/n8n/make-pdf-searchable-or-unsearchable)
* [PDF Merging](/integrations/n8n/merge-pdf)
* [PDF Security](/integrations/n8n/pdf-add-remove-security)
* [Rotate PDF Pages](/integrations/n8n/rotate-pdf)
* [Search and Replace/Delete Text](/integrations/n8n/search-and-replace-text-in-pdf)
* [Search in PDF](/integrations/n8n/search-in-pdf)
* [PDF Splitting](/integrations/n8n/split-pdf)
* [Upload File](/integrations/n8n/upload-file)
* [URL/HTML to PDF Conversion](/integrations/n8n/url-html-to-pdf)
* [Pabbly Connect](/integrations/pabbly-connect)
* [Salesforce](/integrations/salesforce)
* [Sharepoint](/integrations/sharepoint)
* [Add Barcode](/integrations/zapier/add-barcode-to-pdf)
* [Add Form Field](/integrations/zapier/add-formfield-to-pdf)
* [Add Image](/integrations/zapier/add-image-to-pdf)
* [Add Text](/integrations/zapier/add-text-to-pdf)
* [AI Invoice Parser](/integrations/zapier/ai-invoice-parser)
* [Anything to PDF](/integrations/zapier/anything-to-pdf)
* [Barcode Generator](/integrations/zapier/barcode-generator)
* [Advanced Barcode Reader](/integrations/zapier/barcode-reader)
* [Compress & Optimize](/integrations/zapier/compress)
* [Custom API Call](/integrations/zapier/custom-api-call)
* [Document Classifier](/integrations/zapier/document-classifier)
* [Document Parser](/integrations/zapier/document-parser)
* [Send Email With Attachment](/integrations/zapier/email-send)
* [PDF Find Table](/integrations/zapier/find-table)
* [Getting Started with Zapier](/integrations/zapier/getting-started)
* [HTML to PDF](/integrations/zapier/html-to-pdf)
* [Integrating File Sources with PDF.co](/integrations/zapier/input-file-sources)
* [Merge PDF](/integrations/zapier/merge)
* [PDF Filler](/integrations/zapier/pdf-filler)
* [Get PDF Information](/integrations/zapier/pdf-info)
* [PDF Page Tools](/integrations/zapier/pdf-page-tools)
* [PDF Security](/integrations/zapier/pdf-password-and-security)
* [Convert Scanned PDF to Searchable PDF](/integrations/zapier/pdf-searchable)
* [PDF to Anything](/integrations/zapier/pdf-to-anything)
* [Search and Delete Text](/integrations/zapier/search-and-delete-text)
* [Search and Replace Text](/integrations/zapier/search-and-replace-text)
* [Search and Replace With Image](/integrations/zapier/search-and-replace-with-image)
* [Search Text](/integrations/zapier/search-in-pdf)
* [Split PDF Into Multiple Files](/integrations/zapier/split-pdf)
* [Split PDF Based on Barcode Search](/integrations/zapier/split-pdf-by-barcode)
* [Split PDF Based on Text Search](/integrations/zapier/split-pdf-by-text)
## Knowledge Base
* [Overview](/knowledgebase)
* [Convert](/knowledgebase/convert)
* [Create/Edit](/knowledgebase/create-edit)
* [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide)
* [Extract](/knowledgebase/extract)
* [General](/knowledgebase/general)
* [Macros for Text: Built-in and Custom Macros for Auto Text Replacement](/knowledgebase/macros-for-text)
* [Manage](/knowledgebase/manage)
* [Security](/knowledgebase/security)
* [SMTP Configuration Guide](/knowledgebase/smtp-guide)
* [How to setup SMTP for email via AOL mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-aol-mail)
* [How to Set Up SMTP for Email via GMAIL](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-gmail)
* [How to Set Up SMTP for Email via GMX](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-gmx)
* [How to Set Up SMTP for Email via Hushmail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-hushmail)
* [How to Set Up SMTP for Email via iCloud Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-icloud-mail)
* [How to Set Up SMTP for Email via Lycos Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-lycos-mail)
* [How to Set Up SMTP for Email via Mail.com](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-mail-com)
* [How to Set Up SMTP for Email via Office 365 Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-office-365-mail)
* [How to Set Up SMTP for Email via Outlook](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-outlook)
* [How to Set Up SMTP for Email via Postmark](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-postmark)
* [How to Set Up SMTP for Email via Rediffmail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-rediffmail)
* [How to Set Up SMTP for Email via SendGrid](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-sendgrid)
* [How to Set Up SMTP for Email via Verizon](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-verizon)
* [How to Set Up SMTP for Email via Yahoo Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-yahoo-mail)
* [How to Set Up SMTP for Email via Zoho Mail](/knowledgebase/smtp-guide/how-to-setup-smtp-for-email-via-zoho)
* [User-Controlled Encryption](/knowledgebase/user-controlled-encryption)
## API Tester
* [Welcome to PDF.co API Tester](/api-tester)
* [Get Account Balance Info (API Tester)](/api-tester/account-balance-info)
* [AI Invoice Parser (API Tester)](/api-tester/ai-invoice-parser)
* [Barcodes Generator (API Tester)](/api-tester/barcode/generate)
* [Barcodes Reader (API Tester)](/api-tester/barcode/read)
* [Excel to CSV (API Tester)](/api-tester/convert-from-excel/csv)
* [Excel to HTML (API Tester)](/api-tester/convert-from-excel/html)
* [Excel to JSON (API Tester)](/api-tester/convert-from-excel/json)
* [Excel to PDF (API Tester)](/api-tester/convert-from-excel/pdf)
* [Excel to Text (API Tester)](/api-tester/convert-from-excel/text)
* [Excel to XML (API Tester)](/api-tester/convert-from-excel/xml)
* [Document Classifier (API Tester)](/api-tester/document-classifier)
* [Parse Document (API Tester)](/api-tester/documentparser)
* [List All Templates (API Tester)](/api-tester/documentparser/templates)
* [Retrieve Template by ID (API Tester)](/api-tester/documentparser/templates-id)
* [Extract Data from Email File (API Tester)](/api-tester/email/decode)
* [Extract Email Attachment (API Tester)](/api-tester/email/extract-attachments)
* [Send Email with File (API Tester)](/api-tester/email/send)
* [Delete Temporary File (API Tester)](/api-tester/file-upload/delete)
* [Generate Pre-signed URL (API Tester)](/api-tester/file-upload/generate-presigned-url)
* [Get MD5 Hash of File by URL (API Tester)](/api-tester/file-upload/hash)
* [Upload Small File (API Tester)](/api-tester/file-upload/upload)
* [Upload File Using Base64 (API Tester)](/api-tester/file-upload/upload-base64)
* [Upload File from URL (API Tester)](/api-tester/file-upload/upload-url-get)
* [Upload File from URL (API Tester)](/api-tester/file-upload/upload-url-post)
* [PDF Forms Info Reader (API Tester)](/api-tester/forms/info-reader)
* [Background & Job Check (API Tester)](/api-tester/job-check)
* [Merge PDF (API Tester)](/api-tester/merge/pdf)
* [Merge Various Document Type (API Tester)](/api-tester/merge/various-files)
* [PDF Add (API Tester)](/api-tester/pdf-add)
* [Make Text Searchable (API Tester)](/api-tester/pdf-change-text-searchable/searchable)
* [Make Text Unsearchable (API Tester)](/api-tester/pdf-change-text-searchable/unsearchable)
* [PDF Compress (API Tester)](/api-tester/pdf-compress)
* [PDF Delete Pages (API Tester)](/api-tester/pdf-delete-pages)
* [Extract Attachment (API Tester)](/api-tester/pdf-extract-attachments)
* [PDF Find Text (API Tester)](/api-tester/pdf-find/basic)
* [Find Text in Table with AI (API Tester)](/api-tester/pdf-find/table)
* [PDF from CSV (API Tester)](/api-tester/pdf-from-document/csv)
* [PDF from DOC (API Tester)](/api-tester/pdf-from-document/doc)
* [PDF from Email (API Tester)](/api-tester/pdf-from-email)
* [PDF from HTML (API Tester)](/api-tester/pdf-from-html/convert)
* [Return HTML Template by ID (API Tester)](/api-tester/pdf-from-html/template-id)
* [Return All Templates (API Tester)](/api-tester/pdf-from-html/templates)
* [PDF from Image (API Tester)](/api-tester/pdf-from-image)
* [PDF from URL (API Tester)](/api-tester/pdf-from-url)
* [PDF Info Reader (API Tester)](/api-tester/pdf-info-reader)
* [Add Password to PDF (API Tester)](/api-tester/pdf-password/add)
* [Remove Password from PDF (API Tester)](/api-tester/pdf-password/remove)
* [Auto-rotate Pages with AI (API Tester)](/api-tester/pdf-rotate/auto)
* [Rotate Selected Pages (API Tester)](/api-tester/pdf-rotate/basic)
* [PDF Search and Delete Text (API Tester)](/api-tester/pdf-search-text-and-delete)
* [Search and Replace with Image (API Tester)](/api-tester/pdf-search-text-and-replace/image)
* [Search and Replace with Text (API Tester)](/api-tester/pdf-search-text-and-replace/text)
* [Split PDF (API Tester)](/api-tester/pdf-split/by-pages)
* [Split PDF by Text Search (API Tester)](/api-tester/pdf-split/by-text-search-or-barcode)
* [PDF to CSV (API Tester)](/api-tester/pdf-to-csv)
* [PDF to XLS (API Tester)](/api-tester/pdf-to-excel/xls)
* [PDF to XLSX (API Tester)](/api-tester/pdf-to-excel/xlsx)
* [PDF to HTML (API Tester)](/api-tester/pdf-to-html)
* [PDF to JPG (API Tester)](/api-tester/pdf-to-image/jpg)
* [PDF to PNG (API Tester)](/api-tester/pdf-to-image/png)
* [PDF to TIFF (API Tester)](/api-tester/pdf-to-image/tiff)
* [PDF to WEBP (API Tester)](/api-tester/pdf-to-image/webp)
* [PDF to JSON (API Tester)](/api-tester/pdf-to-json/basic)
* [PDF to JSON with AI (API Tester)](/api-tester/pdf-to-json/with-ai)
* [PDF to Text (API Tester)](/api-tester/pdf-to-text/basic)
* [PDF to Text (Simple) (API Tester)](/api-tester/pdf-to-text/simple)
* [PDF to XML (API Tester)](/api-tester/pdf-to-xml)
## Resources
* [Changelog](/changelog)
# Get Account Balance Info (API Tester)
Source: https://developer.pdf.co/api-tester/account-balance-info
/openapi.json get /v1/account/credit/balance
Run Get Account Balance Info live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Get Account Balance Info → API Reference](/api/account-balance-info) — all parameters, response fields, and limits.
# AI Invoice Parser (API Tester)
Source: https://developer.pdf.co/api-tester/ai-invoice-parser
/openapi.json post /v1/ai-invoice-parser
Run AI Invoice Parser live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [AI Invoice Parser → API Reference](/api/ai-invoice-parser) — all parameters, response fields, and limits.
## Prerequisites
Before using the AI Invoice Parser API, please note:
* **Invoices only**: The API processes invoices exclusively to ensure accurate parsing.
* **Asynchronous processing**: When you make a request, you get a JobID immediately while processing happens in the background.
To get your results, you can either:
* Poll the [**job/check**](/api-tester/job-check) endpoint using your `JobID`, or
* Provide a callback URL to get results automatically via webhook.
# Barcodes Generator (API Tester)
Source: https://developer.pdf.co/api-tester/barcode/generate
/openapi.json post /v1/barcode/generate
Run Barcodes Generator live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Barcodes Generator → API Reference](/api/barcode/generate) — all parameters, response fields, and limits.
# Barcodes Reader (API Tester)
Source: https://developer.pdf.co/api-tester/barcode/read
/openapi.json post /v1/barcode/read/from/url
Run Barcodes Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Barcodes Reader → API Reference](/api/barcode/read) — all parameters, response fields, and limits.
# Excel to CSV (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/csv
/openapi.json post /v1/xls/convert/to/csv
Run Excel to CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to CSV → API Reference](/api/convert-from-excel/csv) — all parameters, response fields, and limits.
# Excel to HTML (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/html
/openapi.json post /v1/xls/convert/to/html
Run Excel to HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to HTML → API Reference](/api/convert-from-excel/html) — all parameters, response fields, and limits.
# Excel to JSON (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/json
/openapi.json post /v1/xls/convert/to/json
Run Excel to JSON live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to JSON → API Reference](/api/convert-from-excel/json) — all parameters, response fields, and limits.
# Excel to PDF (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/pdf
/openapi.json post /v1/xls/convert/to/pdf
Run Excel to PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to PDF → API Reference](/api/convert-from-excel/pdf) — all parameters, response fields, and limits.
# Excel to Text (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/text
/openapi.json post /v1/xls/convert/to/txt
Run Excel to Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to Text → API Reference](/api/convert-from-excel/text) — all parameters, response fields, and limits.
# Excel to XML (API Tester)
Source: https://developer.pdf.co/api-tester/convert-from-excel/xml
/openapi.json post /v1/xls/convert/to/xml
Run Excel to XML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Excel to XML → API Reference](/api/convert-from-excel/xml) — all parameters, response fields, and limits.
# Document Classifier (API Tester)
Source: https://developer.pdf.co/api-tester/document-classifier
/openapi.json post /v1/pdf/classifier
Run Document Classifier live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Document Classifier → API Reference](/api/document-classifier) — all parameters, response fields, and limits.
# Parse Document (API Tester)
Source: https://developer.pdf.co/api-tester/documentparser
/openapi.json post /v1/pdf/documentparser
Run Parse Document live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Parse Document → API Reference](/api/documentparser/overview) — all parameters, response fields, and limits.
# List All Templates (API Tester)
Source: https://developer.pdf.co/api-tester/documentparser/templates
/openapi.json get /v1/pdf/documentparser/templates
Run List All Templates live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [List All Templates → API Reference](/api/documentparser/templates) — all parameters, response fields, and limits.
# Retrieve Template by ID (API Tester)
Source: https://developer.pdf.co/api-tester/documentparser/templates-id
/openapi.json get /v1/pdf/documentparser/templates/{id}
Run Retrieve Template by ID live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Retrieve Template by ID → API Reference](/api/documentparser/templates-id) — all parameters, response fields, and limits.
# Extract Data from Email File (API Tester)
Source: https://developer.pdf.co/api-tester/email/decode
/openapi.json post /v1/email/decode
Run Extract Data from Email File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Extract Data from Email File → API Reference](/api/email/decode) — all parameters, response fields, and limits.
# Extract Email Attachment (API Tester)
Source: https://developer.pdf.co/api-tester/email/extract-attachments
/openapi.json post /v1/email/extract-attachments
Run Extract Email Attachment live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Extract Email Attachment → API Reference](/api/email/extract-attachments) — all parameters, response fields, and limits.
# Send Email with File (API Tester)
Source: https://developer.pdf.co/api-tester/email/send
/openapi.json post /v1/email/send
Run Send Email with File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Send Email with File → API Reference](/api/email/send) — all parameters, response fields, and limits.
# Delete Temporary File (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/delete
/openapi.json post /v1/file/delete
Run Delete Temporary File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Delete Temporary File → API Reference](/api/file-upload/delete) — all parameters, response fields, and limits.
# Generate Pre-signed URL (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/generate-presigned-url
/openapi.json get /v1/file/upload/get-presigned-url
Run Generate Pre-signed URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Generate Pre-signed URL → API Reference](/api/file-upload/generate-presigned-url) — all parameters, response fields, and limits.
# Get MD5 Hash of File by URL (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/hash
/openapi.json post /v1/file/hash
Run Get MD5 Hash of File by URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Get MD5 Hash of File by URL → API Reference](/api/file-upload/hash) — all parameters, response fields, and limits.
# Upload Small File (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/upload
/openapi.json post /v1/file/upload
Run Upload Small File live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Upload Small File → API Reference](/api/file-upload/upload) — all parameters, response fields, and limits.
# Upload File Using Base64 (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/upload-base64
/openapi.json post /v1/file/upload/base64
Run Upload File Using Base64 live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Upload File Using Base64 → API Reference](/api/file-upload/upload-base64) — all parameters, response fields, and limits.
# Upload File from URL (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/upload-url-get
/openapi.json get /v1/file/upload/url
Run Upload File from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Upload File from URL → API Reference](/api/file-upload/upload-url-get) — all parameters, response fields, and limits.
# Upload File from URL (API Tester)
Source: https://developer.pdf.co/api-tester/file-upload/upload-url-post
/openapi.json post /v1/file/upload/url
Run Upload File from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Upload File from URL → API Reference](/api/file-upload/upload-url-post) — all parameters, response fields, and limits.
# PDF Forms Info Reader (API Tester)
Source: https://developer.pdf.co/api-tester/forms/info-reader
/openapi.json post /v1/pdf/info/fields
Run PDF Forms Info Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Forms Info Reader → API Reference](/api/forms/info-reader) — all parameters, response fields, and limits.
# Welcome to PDF.co API Tester
Source: https://developer.pdf.co/api-tester/index
The API Tester helps you quickly explore and test our API endpoints. You can send real requests, check responses, and estimate credit usage, all in one place. It's the fastest way to understand how our API works before you integrate it into your own app or workflow.
## Getting Started
To begin testing, you'll need an API Key.
* If you already have an account, [log in](https://app.pdf.co/login) to retrieve your key.
* If not, [sign up](https://app.pdf.co/signup) and receive 10,000 free credits to start exploring the endpoints.
For detailed instructions and examples, see the [Authentication Guide](/api#authenticating-your-api-request).
## URL Input
Our API supports files from any public link, such as Google Drive, Dropbox, and others. However, third-party storage providers may restrict the number of requests to their files, which can cause errors during processing. To prevent this, we recommend using our File Upload endpoint to store files in [PDF.co's built-in storage](https://app.pdf.co/tools/files). For details, see our [File Upload guide](/api/file-upload/overview).
## **Response Codes**
When you run a request, the tester will return both the **output data** and the **response code**. Common codes include:
| Error Code | Description |
| :--------- | :------------------------------------------------------------------------------------------------------------------------------- |
| `200` | Success. |
| `400` | Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password). |
| `401` | Unauthorized. Authentication is required and has failed or has not yet been provided. |
| `402` | Not enough credits. |
| `403` | Access forbidden for input URL. |
| `404` | The requested resource could not be found. |
[**See Full List of Response Codes**](/api/response-codes)
## Credit Usage
Every API call costs credits. The amount depends on the specific endpoint you're using and the size of your file. See [**Credits per API Function**](/api/credits-per-api-function) for the per-endpoint cost, or use our [**Credits Calculator**](https://app.pdf.co/subscriptions#credits-calculator) to estimate usage before sending requests.
## Working with Async Mode
:bulb:Tip: For larger or long-running tasks (over 30 seconds), use async mode to avoid timeouts and optimize credit usage. Set the `async` parameter to true when making your request. The API will return a `JobID` and an empty URL, you can then use the [Job Check endpoint](api-tester/job-check) to retrieve the results once processing is complete. Learn more in the [**Sync and Async Mode Guide**](/api/async-and-sync-mode).
## Test Your Endpoint
You can try out some of the most popular endpoints right here. Simply choose an endpoint, click **Try It**, and provide the file URL you want to process.
Process invoices faster than ever by extracting data and structuring it
automatically with our advanced AI. Get quick and accurate data from any
invoice, no matter the layout.
Add text, images, forms, other PDFs, fill forms, links to external sites and
external PDF files. You can update or modify PDF and scanned PDF files.
Compress PDF files to reduce their size.
Convert PDF and scanned images to text with layout preserved. This method uses
OCR and reporoduces layout.
# Background & Job Check (API Tester)
Source: https://developer.pdf.co/api-tester/job-check
/openapi.json post /v1/job/check
Run Background & Job Check live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Background & Job Check → API Reference](/api/job-check) — all parameters, response fields, and limits.
# Merge PDF (API Tester)
Source: https://developer.pdf.co/api-tester/merge/pdf
/openapi.json post /v1/pdf/merge
Run Merge PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Merge PDF → API Reference](/api/merge/pdf) — all parameters, response fields, and limits.
# Merge Various Document Type (API Tester)
Source: https://developer.pdf.co/api-tester/merge/various-files
/openapi.json post /v1/pdf/merge2
Run Merge Various Document Type live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Merge Various Document Type → API Reference](/api/merge/various-files) — all parameters, response fields, and limits.
# PDF Add (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-add
/openapi.json post /v1/pdf/edit/add
Run PDF Add live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Add → API Reference](/api/pdf-add) — all parameters, response fields, and limits.
# Make Text Searchable (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-change-text-searchable/searchable
/openapi.json post /v1/pdf/makesearchable
Run Make Text Searchable live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Make Text Searchable → API Reference](/api/pdf-change-text-searchable/searchable) — all parameters, response fields, and limits.
# Make Text Unsearchable (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-change-text-searchable/unsearchable
/openapi.json post /v1/pdf/makeunsearchable
Run Make Text Unsearchable live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Make Text Unsearchable → API Reference](/api/pdf-change-text-searchable/unsearchable) — all parameters, response fields, and limits.
# PDF Compress (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-compress
/openapi.json post /v2/pdf/compress
Run PDF Compress live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Compress → API Reference](/api/pdf-compress) — all parameters, response fields, and limits.
# PDF Delete Pages (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-delete-pages
/openapi.json post /v1/pdf/edit/delete-pages
Run PDF Delete Pages live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Delete Pages → API Reference](/api/pdf-delete-pages) — all parameters, response fields, and limits.
# Extract Attachment (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-extract-attachments
/openapi.json post /v1/pdf/attachments/extract
Run Extract Attachment live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Extract Attachment → API Reference](/api/pdf-extract-attachments) — all parameters, response fields, and limits.
# PDF Find Text (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-find/basic
/openapi.json post /v1/pdf/find
Run PDF Find Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Find Text → API Reference](/api/pdf-find/basic) — all parameters, response fields, and limits.
# Find Text in Table with AI (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-find/table
/openapi.json post /v1/pdf/find/table
Run Find Text in Table with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Find Text in Table with AI → API Reference](/api/pdf-find/table) — all parameters, response fields, and limits.
# PDF from CSV (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-document/csv
/openapi.json post /v1/pdf/convert/from/csv
Run PDF from CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from CSV → API Reference](/api/pdf-from-document/csv) — all parameters, response fields, and limits.
# PDF from DOC (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-document/doc
/openapi.json post /v1/pdf/convert/from/doc
Run PDF from DOC live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from DOC → API Reference](/api/pdf-from-document/doc) — all parameters, response fields, and limits.
# PDF from Email (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-email
/openapi.json post /v1/pdf/convert/from/email
Run PDF from Email live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from Email → API Reference](/api/pdf-from-email) — all parameters, response fields, and limits.
# PDF from HTML (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-html/convert
/openapi.json post /v1/pdf/convert/from/html
Run PDF from HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from HTML → API Reference](/api/pdf-from-html/convert) — all parameters, response fields, and limits.
# Return HTML Template by ID (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-html/template-id
/openapi.json get /v1/templates/html/{id}
Run Return HTML Template by ID live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Return HTML Template by ID → API Reference](/api/pdf-from-html/template-id) — all parameters, response fields, and limits.
# Return All Templates (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-html/templates
/openapi.json get /v1/templates/html
Run Return All Templates live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Return All Templates → API Reference](/api/pdf-from-html/templates) — all parameters, response fields, and limits.
# PDF from Image (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-image
/openapi.json post /v1/pdf/convert/from/image
Run PDF from Image live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from Image → API Reference](/api/pdf-from-image) — all parameters, response fields, and limits.
# PDF from URL (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-from-url
/openapi.json post /v1/pdf/convert/from/url
Run PDF from URL live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF from URL → API Reference](/api/pdf-from-url) — all parameters, response fields, and limits.
# PDF Info Reader (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-info-reader
/openapi.json post /v1/pdf/info
Run PDF Info Reader live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Info Reader → API Reference](/api/pdf-info-reader) — all parameters, response fields, and limits.
# Add Password to PDF (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-password/add
/openapi.json post /v1/pdf/security/add
Run Add Password to PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Add Password to PDF → API Reference](/api/pdf-password/add) — all parameters, response fields, and limits.
# Remove Password from PDF (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-password/remove
/openapi.json post /v1/pdf/security/remove
Run Remove Password from PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Remove Password from PDF → API Reference](/api/pdf-password/remove) — all parameters, response fields, and limits.
# Auto-rotate Pages with AI (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-rotate/auto
/openapi.json post /v1/pdf/edit/rotate/auto
Run Auto-rotate Pages with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Auto-rotate Pages with AI → API Reference](/api/pdf-rotate/auto) — all parameters, response fields, and limits.
# Rotate Selected Pages (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-rotate/basic
/openapi.json post /v1/pdf/edit/rotate
Run Rotate Selected Pages live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Rotate Selected Pages → API Reference](/api/pdf-rotate/basic) — all parameters, response fields, and limits.
# PDF Search and Delete Text (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-search-text-and-delete
/openapi.json post /v1/pdf/edit/delete-text
Run PDF Search and Delete Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF Search and Delete Text → API Reference](/api/pdf-search-text-and-delete) — all parameters, response fields, and limits.
# Search and Replace with Image (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-search-text-and-replace/image
/openapi.json post /v1/pdf/edit/replace-text-with-image
Run Search and Replace with Image live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Search and Replace with Image → API Reference](/api/pdf-search-text-and-replace/image) — all parameters, response fields, and limits.
# Search and Replace with Text (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-search-text-and-replace/text
/openapi.json post /v1/pdf/edit/replace-text
Run Search and Replace with Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Search and Replace with Text → API Reference](/api/pdf-search-text-and-replace/text) — all parameters, response fields, and limits.
# Split PDF (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-split/by-pages
/openapi.json post /v1/pdf/split
Run Split PDF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Split PDF → API Reference](/api/pdf-split/by-pages) — all parameters, response fields, and limits.
# Split PDF by Text Search (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-split/by-text-search-or-barcode
/openapi.json post /v1/pdf/split2
Run Split PDF by Text Search live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [Split PDF by Text Search → API Reference](/api/pdf-split/by-text-search-or-barcode) — all parameters, response fields, and limits.
# PDF to CSV (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-csv
/openapi.json post /v1/pdf/convert/to/csv
Run PDF to CSV live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to CSV → API Reference](/api/pdf-to-csv) — all parameters, response fields, and limits.
# PDF to XLS (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-excel/xls
/openapi.json post /v1/pdf/convert/to/xls
Run PDF to XLS live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to XLS → API Reference](/api/pdf-to-excel/xls) — all parameters, response fields, and limits.
# PDF to XLSX (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-excel/xlsx
/openapi.json post /v1/pdf/convert/to/xlsx
Run PDF to XLSX live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to XLSX → API Reference](/api/pdf-to-excel/xlsx) — all parameters, response fields, and limits.
# PDF to HTML (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-html
/openapi.json post /v1/pdf/convert/to/html
Run PDF to HTML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to HTML → API Reference](/api/pdf-to-html) — all parameters, response fields, and limits.
# PDF to JPG (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-image/jpg
/openapi.json post /v1/pdf/convert/to/jpg
Run PDF to JPG live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to JPG → API Reference](/api/pdf-to-image/jpg) — all parameters, response fields, and limits.
# PDF to PNG (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-image/png
/openapi.json post /v1/pdf/convert/to/png
Run PDF to PNG live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to PNG → API Reference](/api/pdf-to-image/png) — all parameters, response fields, and limits.
# PDF to TIFF (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-image/tiff
/openapi.json post /v1/pdf/convert/to/tiff
Run PDF to TIFF live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to TIFF → API Reference](/api/pdf-to-image/tiff) — all parameters, response fields, and limits.
# PDF to WEBP (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-image/webp
/openapi.json post /v1/pdf/convert/to/webp
Run PDF to WEBP live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to WEBP → API Reference](/api/pdf-to-image/webp) — all parameters, response fields, and limits.
# PDF to JSON (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-json/basic
/openapi.json post /v1/pdf/convert/to/json2
Run PDF to JSON live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to JSON → API Reference](/api/pdf-to-json/basic) — all parameters, response fields, and limits.
# PDF to JSON with AI (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-json/with-ai
/openapi.json post /v1/pdf/convert/to/json-meta
Run PDF to JSON with AI live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to JSON with AI → API Reference](/api/pdf-to-json/with-ai) — all parameters, response fields, and limits.
# PDF to Text (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-text/basic
/openapi.json post /v1/pdf/convert/to/text
Run PDF to Text live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to Text → API Reference](/api/pdf-to-text/basic) — all parameters, response fields, and limits.
# PDF to Text (Simple) (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-text/simple
/openapi.json post /v1/pdf/convert/to/text-simple
Run PDF to Text (Simple) live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to Text (Simple) → API Reference](/api/pdf-to-text/simple) — all parameters, response fields, and limits.
# PDF to XML (API Tester)
Source: https://developer.pdf.co/api-tester/pdf-to-xml
/openapi.json post /v1/pdf/convert/to/xml
Run PDF to XML live from your browser — send a real request and see the response and credit usage, no code required. PDF.co API Tester.
**Full reference:** [PDF to XML → API Reference](/api/pdf-to-xml) — all parameters, response fields, and limits.
# Get Account Balance Info
Source: https://developer.pdf.co/api/account-balance-info
Get remaining account balance.
**Try it live:** [Get Account Balance Info → API Tester](/api-tester/account-balance-info) — send a real request from your browser.
## `GET /v1/account/credit/balance`
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------ |
| `remainingCredits` | integer | Number of credits remaining in the account |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/account/credit/balance' \
--header 'x-api-key: '
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"remainingCredits": 99795868
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/account/credit/balance' \
--header 'x-api-key: '
```
# AI Invoice Parser
Source: https://developer.pdf.co/api/ai-invoice-parser
Extract structured data from invoices of any layout using AI-based parsing, without requiring per-vendor templates.
**Try it live:** [AI Invoice Parser → API Tester](/api-tester/ai-invoice-parser) — send a real request from your browser.
## `POST /ai-invoice-parser`
The AI Invoice parser automatically detects invoice layouts without the manual effort previously required to supply document parsing templates for reference.
**Important**
* **Only invoices will be parsed**. For all other documents, please use our existing [**Document Parser**](/api/documentparser/parser).
* To ensure accurate processing, each invoice must be clearly separated. **If an invoice contains multiple pages, we recommend splitting it** into individual PDFs using the [PDF Split API](/api/pdf-split/by-pages).
* While AI Invoice Parser supports multi-page invoices, **the total page count for a single PDF must not exceed 100 pages**. Submitting large PDFs containing multiple invoices is not recommended.
* To retrieve results, you must poll the [Background Job Check](/api/job-check) endpoint using the `jobId` returned in the initial response. Once the job status is marked as `success`, the output file will be available at the provided URL.
This method extracts data from your PDF invoices and returns a [well-structured JSON format](/api/ai-invoice-parser/#example-response) for your use.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------- | ------ | -------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url attribute`](/api/url-input-and-request-limits#supported-file-sources) |
| `customField` | string | *No* | - | JSON string containing [custom field](/api/ai-invoice-parser/#custom-fields) names to extract. Use `camelCase` for field names (e.g., `storeNumber`, `deliveryDate`). Multiple fields should be comma-separated. |
| `lineItemStructure` | object | *No* | - | Defines a custom structure for line items in the response. See [Line Item Structure](/api/ai-invoice-parser/#line-item-structure) for more information. |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Custom Fields
AI Invoice Parser with custom fields support automatically detects invoice layouts and extracts both standard schema data and user-specified custom fields without requiring manual templates.
The `customField` parameter allows you to specify additional fields to extract beyond the standard schema. Some examples include:
* `storeNumber` - Store or branch identifier
* `deliveryDate` - Expected delivery date
* `financialCharges` - Additional financial charges
* `lineTotal` - Total amount for line items
* `purchaseOrderRef` - Purchase order reference number
* `customerReference` - Customer reference number
* `departmentCode` - Department or cost center code
If a custom field returns an empty value, please [contact our support team](https://pdf.co/support/request?subject=ai-invoice-parser%20-%20custom%20fields) to help improve the extraction accuracy.
## Line Item Structure
The `lineItemStructure` attribute lets you define a custom schema for line items. Each key is a field name you choose (in `camelCase`) and each value is the expected data type — either `"string"` or `"number"`.
When provided, every object inside the `lineItems` array will contain exactly the fields you specified. If a value cannot be extracted from the invoice, the field will still be present with an empty or default value instead of being omitted.
### Example
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"async": true,
"lineItemStructure": {
"description": "string",
"quantity": "number",
"unitPrice": "number",
"totalPrice": "number"
}
}
```
With the structure above, every line item in the response will include all four fields:
```json theme={null}
"lineItems": [
[
{
"description": "Item 1",
"quantity": 2,
"unitPrice": 9.95,
"totalPrice": 19.90
},
{
"description": "Item 2",
"quantity": 5,
"unitPrice": 20.00,
"totalPrice": 100.00
}
]
]
```
Use `camelCase` for field names (e.g., `unitPrice`, `totalPrice`). The field names you define will be used as-is in the response, giving you full control over the output keys.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | -------------------------------------------------------------------------------------- |
| `status` | string | Status of the API response. The statuses are: `success`, `error`. |
| `message` | string | Descriptive message for the response status. |
| `pageCount` | integer | Number of pages processed or returned. |
| `body` | object | Contains the invoice data. See [Invoice Schema](#invoice-schema) for more information. |
| `jobId` | string | Unique identifier for the background job. |
| `credits` | integer | Credits used for this operation. |
| `remainingCredits` | integer | Credits left after this job execution. |
| `duration` | integer | Time taken to complete the request, in milliseconds. |
### Invoice Schema
The `body` object contains all the metadata needed to understand your invoice content and includes the following attributes:
```json theme={null}
"body": {
"vendor": { .... },
"customer": { .... },
"invoice": { .... },
"paymentDetails": { .... },
"others": { .... },
"lineItems": { .... }
}
```
#### Sections
* [vendor](#vendor-object)
* [customer](#customer-object)
* [invoice](#invoice-object)
* [paymentDetails](#paymentdetails-object)
* [others](#others-object)
* [lineItems](#lineitems-object)
***
### The `vendor` Object
An `object` containing vendor details.
| Attribute | Type | Description |
| -------------------- | ------ | ----------------------------------------------------------------------------------- |
| `name` | string | Name of the vendor |
| `address` | object | Vendor's address details. See [address object](#address-object) |
| `contactInformation` | object | Vendor contact details. See [contactInformation object](#contactinformation-object) |
| `entityId` | object | Vendor's entity ID (e.g., EIN, ABN, VAT, GST, etc.) |
***
### The `customer` Object
An `object` containing customer details.
| Attribute | Type | Description |
| --------- | ------ | -------------------------------------------------------- |
| `billTo` | object | Billing details. See [customer.billTo](#customerbillto) |
| `shipTo` | object | Shipping details. See [customer.shipTo](#customershipto) |
### `customer.billTo`
| Attribute | Type | Description |
| -------------------- | ------ | ----------------------------------------------------------- |
| `name` | string | Customer name |
| `address` | object | See [address object](#address-object) |
| `contactInformation` | object | See [contactInformation object](#contactinformation-object) |
| `entityId` | string | Customer's entity ID (e.g., EIN, ABN, VAT, GST, etc.) |
### `customer.shipTo`
| Attribute | Type | Description |
| --------- | ------ | ------------------------------------- |
| `name` | string | Customer name |
| `address` | object | See [address object](#address-object) |
***
### The `invoice` Object
An `object` containing the invoice details.
| Attribute | Type | Description |
| ------------- | ------ | --------------------- |
| `invoiceNo` | string | Invoice number |
| `invoiceDate` | string | Date of invoice |
| `poNo` | string | Purchase order number |
| `orderNo` | string | Sales order number |
***
### The `paymentDetails` Object
An `object` containing payment details.
| Attribute | Type | Description |
| -------------------- | ------ | ----------------------------------------------------------- |
| `paymentTerms` | string | Terms of payment |
| `dueDate` | string | Payment due date |
| `total` | string | Total amount due |
| `subtotal` | string | Subtotal amount |
| `tax` | string | Tax amount |
| `discount` | string | Discount amount |
| `shipping` | string | Shipping amount |
| `bankingInformation` | object | See [bankingInformation](#paymentdetailsbankinginformation) |
### `paymentDetails.bankingInformation`
| Attribute | Type | Description |
| ------------------- | ------ | -------------------------------------------------- |
| `bankName` | string | Name of the bank |
| `accountHolderName` | string | Name of the account holder |
| `accountNumber` | string | Bank account number |
| `iban` | string | International Bank Account Number (IBAN) |
| `swiftBicCode` | string | SWIFT/BIC code of the bank |
| `bankAddress` | object | See [address object](#address-object) |
| `bankRoutingCode` | string | Routing code for domestic payments |
| `bankCode` | string | Institution number within Canadian banking network |
| `branchNumber` | string | Branch-specific code |
| `purposeCode` | string | Specifies the transaction's intent |
| `additionalNotes` | string | Payment instructions or other notes |
***
### The `others` Object
An `object` containing additional notes.
| Attribute | Type | Description |
| --------- | ------ | ---------------------------------------------- |
| `notes` | string | Additional notes such as delivery instructions |
***
### The `lineItems` Object
An `object` detailing the line items in an invoice.
Note: there is no common structure due to significant variability between invoices! To define your own custom structure, use the [`lineItemStructure`](/api/ai-invoice-parser/#line-item-structure) attribute in your request.
A typical invoice might list purchase items with details such as name, quantity or price of each individual item.
***
### Common Objects
There are a couble of objects which are commonly used in the schema in a few places, these are as follows.
### `address` object
| Attribute | Type | Description |
| --------------- | ------ | ----------------- |
| `streetAddress` | string | Street address |
| `city` | string | City name |
| `state` | string | State/county name |
| `postalCode` | string | Postal/ZIP code |
| `country` | string | Country code/name |
### `contactInformation` object
| Attribute | Type | Description |
| --------- | ------ | --------------------------- |
| `phone` | string | Phone number of the vendor |
| `fax` | string | Fax number of the vendor |
| `email` | string | Email address of the vendor |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"callback": "https://example.com/callback/url/you/provided"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes).
You can use jobId to identify the corresponding callback response. Use the [Job Check](/api/job-check) API to poll the job status.
```json theme={null}
{
"error": false,
"status": "created",
"jobId": "7830deca-2e66-11ef-9ad3-8eff830e7461",
"credits": 100,
"remainingCredits": 106674,
"duration": 33
}
```
## `Example` Callback Response
```json theme={null}
{
"status": "success",
"message": "Success",
"pageCount": 1,
"body": {
"vendor": {
"name": "ACME Inc.",
"address": {
"streetAddress": "1540 Long Street",
"city": "Jacksonville",
"state": "FL",
"postalCode": "32099",
"country": "US"
},
"contactInformation": {
"phone": "352-200-0371",
"fax": "904-787-9468"
}
},
"customer": {
"billTo": {
"name": "Lanny Lane Ltd.",
"address": {
"streetAddress": "82 Gorby Lane",
"city": "Columbia",
"state": "IN",
"postalCode": "39429",
"country": "US"
}
},
"shipTo": {
"name": "Same as recipient"
}
},
"invoice": {
"invoiceNo": "67893566",
"invoiceDate": "JAN 5, 2025"
},
"paymentDetails": {
"total": "$1,272.35",
"subtotal": "$1,262.35",
"tax": "$10.00",
"shipping": "$0.00"
},
"lineItems": [
[
{
"quantity": "2",
"description": "Item 1",
"unit_price": "9.95",
"total": "19.90"
},
{
"quantity": "5",
"description": "Item 2",
"unit_price": "20.00",
"total": "100.00"
},
{
"quantity": "1",
"description": "Item 3",
"unit_price": "19.95",
"total": "19.95"
},
{
"quantity": "1",
"description": "Item 4",
"unit_price": "123.00",
"total": "123.00"
},
{
"quantity": "10",
"description": "Item 5",
"unit_price": "99.95",
"total": "999.50"
}
]
]
},
"jobId": "7830deca-2e66-11ef-9ad3-8eff830e7461",
"credits": 100,
"remainingCredits": 106472,
"duration": 33
}
```
## Setting up the Callback URL
The callback URL should be a webhook which listens to responses from the parsing results. You can setup your own webhook or use one from a provider.
If you are unsure about webhooks or callbacks, please read this [Wikipedia article](https://en.wikipedia.org/wiki/Webhook) to get started.
## Supported Languages
* **Albanian (Shqip)**
* **Bosnian (Bosanski)**
* **Bulgarian (Български)**
* **Croatian (Hrvatski)**
* **Czech (Čeština)**
* **Danish (Dansk)**
* **Dutch (Nederlands)**
* **English**
* **Estonian (Eesti)**
* **Finnish (Suomi)**
* **French (Français)**
* **German (Deutsch)**
* **Greek (Ελληνικά)**
* **Hungarian (Magyar)**
* **Icelandic (Íslenska)**
* **Italian (Italiano)**
* **Latvian (Latviešu)**
* **Lithuanian (Lietuvių)**
* **Norwegian (Norsk)**
* **Polish (Polski)**
* **Portuguese (Português)**
* **Romanian (Română)**
* **Russian (Русский)**
* **Serbian (Српски)**
* **Slovak (Slovenčina)**
* **Slovenian (Slovenščina)**
* **Spanish (Español)**
* **Swedish (Svenska)**
* **Turkish (Türkçe)**
* **Ukrainian (Українська)**
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl -X POST \
https://api.pdf.co/v1/ai-invoice-parser
```
```javascript theme={null}
var https = require("https");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "YOUR_API_KEY_HERE";
// Direct URL of the source PDF file
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf";
// Prepare request to `AI Invoice Parser` API endpoint
var queryPath = `/v1/ai-invoice-parser`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
var postRequest = https.request(reqOptions, (response) => {
let responseData = '';
response.on("data", (chunk) => {
responseData += chunk;
});
response.on("end", () => {
try {
// Parse JSON response
var data = JSON.parse(responseData);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log(data.message);
}
} catch (error) {
console.error("Error parsing JSON response:", error);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
function checkIfJobIsCompleted(jobId, resultFileUrl) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
let responseData = '';
response.setEncoding("utf8");
response.on("data", (chunk) => {
responseData += chunk;
});
response.on("end", () => {
try {
// Parse JSON response
let data = JSON.parse(responseData);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000);
}
else if (data.status == "success") {
console.log("** Response **")
console.log(data);
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
} catch (error) {
console.error("Error parsing JSON response:", error);
}
});
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"
def main(args = None):
getParsedInvoice(SourceFileURL)
def getParsedInvoice(uploadedFileUrl):
"""AI Invoice Parser using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {}
parameters["url"] = uploadedFileUrl
# Prepare URL for 'AI Invoice Parser' API request
url = "{}/ai-invoice-parser".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if(json["status"]):
print("** Response **")
print(json)
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.Net;
using System.Threading;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of Source PDF file
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `AI Invoice Parser` API call
string url = "https://api.pdf.co/v1/ai-invoice-parser";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("url", SourceFileURL);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Asynchronous job ID
string jobId = json["jobId"].ToString();
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do
{
string job_response = "";
string status = CheckJobStatus(jobId, out job_response); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status);
if (status == "success")
{
Console.WriteLine("** Final Response **");
Console.WriteLine(job_response);
break;
}
else if (status == "working")
{
// Pause for a few seconds
Thread.Sleep(3000);
}
else
{
Console.WriteLine(status);
break;
}
}
while (true);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
static string CheckJobStatus(string jobId, out string response)
{
using (WebClient webClient = new WebClient())
{
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
response = webClient.DownloadString(url);
JObject json = JObject.Parse(response);
return Convert.ToString(json["status"]);
}
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.time.LocalDateTime;
import java.time.format.DateTimeFormatter;
public class Main {
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "********************************";
// (!) Make asynchronous job
final static boolean Async = true;
public static void main(String[] args) throws IOException {
// Source PDF file
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf";
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// AI PARSE INVOICE
ParseInvoice(webClient, SourceFileUrl);
}
public static void ParseInvoice(OkHttpClient webClient, String uploadedFileUrl) throws IOException {
// Prepare POST request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("url", new JsonPrimitive(uploadedFileUrl));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare URL for AI Invoice Parser API call.
// See documentation: https://developer.pdf.co/api/ai-invoice-parser
String query = "https://api.pdf.co/v1/ai-invoice-parser";
DateTimeFormatter dtf = DateTimeFormatter.ofPattern("MM/dd/yyyy HH:mm:ss");
// Prepare request to `Document Parser` API
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200) {
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error) {
// Asynchronous job ID
String jobId = json.get("jobId").getAsString();
System.out.println("Job#" + jobId + ": has been created. - " + dtf.format(LocalDateTime.now()));
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do {
String status = CheckJobStatus(webClient, jobId); // Possible statuses: "working", "failed", "aborted", "success"
System.out.println("Job#" + jobId + ": " + status + " - " + dtf.format(LocalDateTime.now()));
if (status.compareToIgnoreCase("success") == 0) {
break;
} else if (status.compareToIgnoreCase("working") == 0) {
// Pause for a few seconds
try {
Thread.sleep(3000);
} catch (InterruptedException ex) {
Thread.currentThread().interrupt(); // restore interrupted status
}
} else {
System.out.println(status);
break;
}
} while (true);
} else {
// Display service reported error
System.out.println(json.get("message").getAsString());
}
} else {
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
// Check Job Status
private static String CheckJobStatus(OkHttpClient webClient, String jobId) throws IOException {
String url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
String status = "";
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200) {
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
status = json.get("status").getAsString();
if(status.equals("success")){
System.out.println(json);
}
return status;
} else {
// Display request error
System.out.println(response.code() + " " + response.message());
}
return "Failed";
}
}
```
```php theme={null}
AI Invoice Parser example.
";
if (curl_errno($curl) == 0)
{
$status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE);
if ($status_code == 200)
{
$json = json_decode($result, true);
if (!isset($json["error"]) || $json["error"] == false)
{
// Asynchronous job ID
$jobId = $json["jobId"];
// Check the job status in a loop
do
{
$status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
echo "
" . date(DATE_RFC2822) . ": " . $status . "
";
if ($status == "success")
{
break;
}
else if ($status == "working")
{
// Pause for a few seconds
sleep(3);
}
else
{
echo $status . " ";
break;
}
}
while (true);
}
else
{
// Display service reported error
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Understanding Sync and Async Modes
Source: https://developer.pdf.co/api/async-and-sync-mode
Compare Sync and Async request modes, including time limits, jobId-based status checks, and credit cost differences.
When you use APIs, the mode you choose Synchronous (Sync) or Asynchronous (Async) can significantly affect both performance and cost. This article explains the differences between these modes, highlights why Async is superior, and shows how switching can benefit you.
## What are Sync and Async Modes?
### Synchronous (Sync) Mode
In **Sync** mode, your API request is processed immediately, and the result is returned within a limit of 30 seconds. Here are a few disadvantages:
* **API Response**: You receive your results when the processing is complete.
* **No Job ID**: There’s no `jobId` provided by [job/check](/api/job-check) for further status checks.
* **Time Limit**: If processing takes longer than 30 seconds, your request will fail.
* **Higher Costs**: Credits are calculated as:
```javascript theme={null}
Total Credits Used = Endpoint Credits × Number of Pages
```
For precise credit calculations, see [Credits per API Function](/api/credits-per-api-function) for the per-endpoint cost, or use the interactive [Credits Calculator](https://app.pdf.co/subscriptions#credits-calculator).
### Asynchronous (Async) Mode
To retrieve results, you must poll the [Background Job Check](/api/job-check) endpoint using the `jobId` returned in the initial response. Once the job status is marked as `success`, the output file will be available at the provided URL.
**Async** mode processes your request in the background, providing the option to track progress which gives you:
* **Immediate API Response**: You receive a `jobId` and an output URL right away.
* **Background Processing**: Handles larger tasks without immediate timeouts.
* **Extended Time Limit**: Can process tasks for up to 3 minutes, reducing timeouts.
* **Status Checks**: Use the `jobId` to monitor progress via the [job/check](/api/job-check) endpoint.
* **Webhook Support**: Supports callback to notify you when your job is done on a specified webhook URL.
* **Lower Cost**: Credits are calculated as:
```javascript theme={null}
Total Credits Used = Endpoint Credits + ("Job/Check" Credits × Number of "Job/Check" Calls until status is "working")
```
## Why Async Mode is Superior
1. **Handles Bigger Tasks Efficiently**: Designed for larger files or complex tasks that exceed Sync mode’s 30-second limit, reducing failures and saving time.
2. **More Cost-Effective**: Although Async mode includes a small cost for each [job/check](/api/job-check) call (credits charged per check until status is “working”), it often results in overall savings by minimizing failed requests and unnecessary retries.
3. **Better Control and Transparency**: The `jobId` allows you to check your job’s status, giving you more control and clear insight into the process.
4. **Fewer Timeouts and Failures**: Extended time limits decrease the likelihood of failures.
5. **Automates Workflow with Webhooks**: [Webhooks](/api/webhooks) notify you automatically when your job is complete, reducing the need for manual checks and further saving on [job/check](/api/job-check) credits.
## How to Switch to Async Mode
1. **Modify Your API Request**:
* Specify that you want to use Async mode in your API call, often by setting an `async` parameter to `true`.
2. **Receive the** `jobId` **and Output URL**:
* After submitting your request, you’ll get a `jobId` and an output URL where results will be available once processing is complete.
3. **Implement Status Checks (Optional)**:
* Use the [job/check](/api/job-check) endpoint with your `jobId` to monitor progress.
* Remember, each [job/check](/api/job-check) call costs credits until status is “working”.
4. **Set Up Webhooks (Recommended)**:
* Configure [Webhooks & Callbacks](/api/webhooks) to receive automatic notifications when your job is finished (costs 2 credits).
* This reduces the need for manual status checks and saves on [job/check](/api/job-check) credits.
## Conclusion
Async mode offers a superior approach for API interactions by providing efficiency, cost-effectiveness, and enhanced control. By switching from Sync to Async mode, you optimize resource usage and unlock features like webhooks and extended processing times.
## Addressing Common Concerns
### What About the Cost of Multiple Status Checks?
Async mode is designed to reduce the need for frequent status checks. By setting up [webhooks](/api/webhooks), you receive automatic updates when your job is complete, which minimizes the number of [job/check](/api/job-check) calls and optimizes credit usage.
### Is Switching to Async Mode Complicated?
No, it’s straightforward. You handle the `jobId` and output URL provided after your initial request. Implementing [webhooks](/api/webhooks) can make the process even smoother by automating notifications.
# Barcodes Generator
Source: https://developer.pdf.co/api/barcode/generate
Generate high quality barcode images. Supports QR Code, Datamatrix, Code 39, Code 128, PDF417 and many other barcode types.
**Try it live:** [Barcodes Generator → API Tester](/api-tester/barcode/generate) — send a real request from your browser.
## `POST /v1/barcode/generate`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `type` | string | *Yes* | QRCode | Set the barcode type to be used. See available barcode types in the [Supported Barcode Types](/api/barcode/overview#supported-barcode-types) |
| `value` | string | *Yes* | - | Set the string value to encode inside the barcode, must be in a string format. |
| `decorationImage` | string | *No* | - | Set this to the image that you want to be inserted the logo inside the QR-Code barcode. To use your file please upload it first to the temporary storage, see the [Upload Files](/api/file-upload/overview) section. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `Angle` | integer | *No* | `0` | See [profiles.Angle](#profiles-angle) |
| `NarrowBarWidth` | integer | *No* | 3 | See [profiles.NarrowBarWidth](#profiles-narrowbarwidth) |
| `CaptionFont` | string | *No* | Arial, 12 | See [profiles.CaptionFont](#profiles-captionfont) |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
#### `profiles.Angle`
Specifies the barcode’s rotation angle as an integer in degrees.
| Value | Description |
| ----- | --------------------- |
| 0 | 0 degrees clockwise |
| 1 | 90 degrees clockwise |
| 2 | 180 degrees clockwise |
| 3 | 270 degrees clockwise |
```
{
"profiles": "{'Angle': 3}"
}
```
#### `profiles.NarrowBarWidth`
Specifies the width of the narrow bars in the barcode in pixels.
```
{
"profiles": "{'NarrowBarWidth': 3}"
}
```
#### `profiles.CaptionFont`
Specifies the font and size of the caption text displayed with the barcode.
```
{
"profiles": "{'CaptionFont': 'Arial, 12'}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
### QRCode Example
#
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": false,
"async": false,
"decorationImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-generator/logo.png"
}
```
### `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/72bc579b37844d9f9e63ce06de5196d8/barcode.png",
"error": false,
"status": 200,
"name": "barcode.png",
"duration": 380,
"remainingCredits": 98725598,
"credits": 7
}
```
#### `Example` CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": false,
"async": false
}'
```
### QRCode with Logo Inside Example
#
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": true,
"async": false
}
```
### `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/9a87556a8b9e4f4eae60843e697250d4/barcode.png",
"error": false,
"status": 200,
"name": "barcode.png",
"remainingCredits": 60631
}
```
#### `Example` CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": false,
"async": false,
"decorationImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-generator/logo.png"
}'
```
### Data URI as Output Example
#
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": false,
"async": false
}
```
### `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAEsAAABLCAYAAAA4TnrqAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAAAOlSURBVHhe7ZBBimMxEMVy/0v34CELkSmBjH96NhaIwKtXlY9fP5fMfawN7mNtcB9rg/tYG9zH2kAf6/V6Pa5hHeaUTPNTjftYg8Z9rEEjPdYJdoc5JZaT0imUOzopywW7w5wSy0npFModnZTlgt1hTonlpHQK5Y5ObJm5SUpODeuU3CSWE53YMnOTlJwa1im5SSwnOrFl5iYpOTWsU3KTWE50YsvMTWI5sY7lxDrMTWI50YktMzeJ5cQ6lhPrMDeJ5UQntszcJJYT61hOrMPcJJYTndgyc5Ps5ob1S24Sy4lObJm5SXZzw/olN4nlRCe2zNwku7lh/ZKbxHKik7JcsDuWE3YosXyXckcnZblgdywn7FBi+S7ljk7KcsHuWE7YocTyXcodnXD5Kck38qc0dDIdOZV8I39KQyfTkVPJN/KnNHzyZaaPrP4v7mNtcB9rA/3n6SOXxHLCDiXTfFmY9j4l03xZ0NZ0cEksJ+xQMs2XhWnvUzLNlwVtTQeXxHLCDiXTfFmY9j4l03xZSK3p+JJYTtgxC9Pe0rAOc2qkr5sOLonlhB2zMO0tDeswp0b6uungklhO2DEL097SsA5zaqSvs0PMi8Zuxzyh3En/YIeYF43djnlCuZP+wQ4xLxq7HfOEcmf7H+yo5WS3Q42puySWk9R5/2bsqOVkt0ONqbsklpPUef9m7KjlZLdDjam7JJaT1Hn/fg1+hElKTo3SIaXfLh3AjzBJyalROqT026UD+BEmKTk1SoeUfrv0BdLHHXSYUyN13r+/Tvq4gw5zaqTO+/fXSR930GFOjdR5//4Dl5+SWF7gLiWWk9Ih2uKhpySWF7hLieWkdIi2eOgpieUF7lJiOSkdoq3dQ8bJHe5SY+oujam7NHRSlgsnd7hLjam7NKbu0tBJWS6c3OEuNabu0pi6S0MntszcJJYb7NPCbp+UXZ3YMnOTWG6wTwu7fVJ2dWLLzE1iucE+Lez2SdnViS0zN0nJTWPqVsk0Xxo6sWXmJim5aUzdKpnmS0MntszcJCU3jalbJdN8aejElpmbxPJdyh12KNnNiU5smblJLN+l3GGHkt2c6MSWmZvE8l3KHXYo2c2JTspywe4wp8TywsmuoZee+jO7w5wSywsnu4ZeeurP7A5zSiwvnOwaeol/9pTEcsIONabup8RyQ1s89JTEcsIONabup8RyQ1s89JTEcsIONabup8Ryo7Uuf7mPtcF9rA3uY21wH2uD+1iZn58/9whzEbhRquEAAAAASUVORK5CYII=",
"error": false,
"status": 200,
"name": "barcode.png",
"duration": 298,
"remainingCredits": 98725605,
"credits": 7
}
```
#### `Example` CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/barcode/generate' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"name": "barcode.png",
"value": "abcdef123456",
"type": "QRCode",
"inline": true,
"async": false
}'
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```javascript theme={null}
var https = require("https");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Result image file name
const DestinationFile = "./barcode.png";
// Barcode type. See valid barcode types in the documentation https://developer.pdf.co
const BarcodeType = "Code128";
// Barcode value
const BarcodeValue = "qweasd123456";
// Prepare request to `Barcode Generator` API endpoint
var queryPath = `/v1/barcode/generate`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: 'barcode.png',
type: BarcodeType,
value: BarcodeValue
});
var reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
exports.handler = async (event) => {
let dataString = '';
const promise_response = await new Promise((resolve, reject) => {
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on('data', chunk => {
dataString += chunk;
});
response.on('end', () => {
resolve({
statusCode: 200,
body: JSON.stringify(JSON.parse(dataString), null, 4)
});
});
}).on("error", (e) => {
reject({
statusCode: 500,
body: 'Something went wrong!'
});
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
});
return promise_response;
};
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Result file name
ResultFile = ".\\barcode.png"
# Barcode type. See valid barcode types in the documentation https://developer.pdf.co
BarcodeType = "Code128"
# Barcode value
BarcodeValue = "qweasd123456"
def main(args = None):
generateBarcode(ResultFile)
def generateBarcode(destinationFile):
"""Generates Barcode using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/barcode/generate
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["type"] = BarcodeType
parameters["value"] = BarcodeValue
# Prepare URL for 'Barcode Generate' API request
url = "{}/barcode/generate".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFCOWebApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Result file name
const string ResultFileName = @".\barcode.png";
// Barcode type. See valid barcode types in the documentation https://developer.pdf.co
const string BarcodeType = "Code128";
// Barcode value
const string BarcodeValue = "qweasd123456";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare requests params as JSON
// See documentation: https://developer.pdf.co/#barcode-generator
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(ResultFileName));
parameters.Add("type", BarcodeType);
parameters.Add("value", BarcodeValue);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// URL of "Barcode Generator" endpoint
string url = "https://api.pdf.co/v1/barcode/generate";
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated barcode image file
string resultFileURI = json["url"].ToString();
// Download generated image file
webClient.DownloadFile(resultFileURI, ResultFileName);
Console.WriteLine("Generated barcode saved to \"{0}\" file.", ResultFileName);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
finally
{
webClient.Dispose();
}
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Result file name
final static Path ResultFile = Paths.get(".\\barcode.png");
// Barcode type. See valid barcode types in the documentation https://developer.pdf.co
final static String BarcodeType = "Code128";
// Barcode value
final static String BarcodeValue = "qweasd123456";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Barcode Generator` API call
String query = "https://api.pdf.co/v1/barcode/generate";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"type\": \"%s\", \"value\": \"%s\"}",
ResultFile.getFileName(),
BarcodeType,
BarcodeValue);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated barcode image file
String resultFileUrl = json.get("url").getAsString();
// Download the image file
downloadFile(webClient, resultFileUrl, ResultFile);
System.out.printf("Generated barcode saved to \"%s\" file.", ResultFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile.toFile());
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
## Result:";
}
else
{
// Display service reported errors
echo "
Error: " . $json["message"] . "
";
}
}
else
{
// Display request error
echo "
Status code: " . $status_code . "
";
echo "
" . $result . "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
```
# Barcodes Overview
Source: https://developer.pdf.co/api/barcode/overview
Generate and read barcodes.
## Supported Barcode Types
| Name | Type | Character Set | Length | Notes |
| --------------------------- | ------- | ------------------------------------------------------------------------------------- | ---------------------------------------- | ------------------------------------------------------------------- |
| `AustralianPostCode` | 2D | Numbers Only | 4 | - |
| `Aztec` | 2D | Full ASCII; FNC1 and ESI control codes | Variable, Min 12 - Max 3832 | - |
| `Codabar` | Linear | Numbers: 0-9; Symbols: - : . \$ / + Start/Stop Characters: A, B, C, D, E, \*, N, or T | Variable | - |
| `CodablockF` | Complex | - | - | See this guide. |
| `Code128` | Linear | All ASCII characters and control codes | Variable | - |
| `Code16K` | - | - | - | - |
| `Code39` | Linear | Uppercase letters A-Z; Numbers 0-9; Space - . \$ / + % | Variable | - |
| `Code39Extended` | Linear | All ASCII characters and control codes | Variable | - |
| `Code39Mod43` | - | - | - | - |
| `Code39Mod43Extended` | - | - | - | - |
| `Code93` | Linear | Uppercase letters A-Z; Numbers 0-9; Space - . \$ / + % | - | - |
| `DataMatrix` | 2D | All ASCII characters | Variable | - |
| `DPMDataMatrix` | - | - | - | - |
| `EAN13` | Linear | Numbers Only | 13 + check digit +2 optional +5 optional | - |
| `EAN2` | Linear | Numbers Only | Exact 2 Numbers | - |
| `EAN5` | Linear | Numbers Only | Exact 5 Numbers | - |
| `EAN8` | Linear | Numbers Only | 7 + check digit | - |
| `GS1 - 128` | Linear | ASCII symbols | 128 ASCII symbols | - |
| `GS1DataBarExpanded` | Linear | String | 74 numeric or 41 alphabetic characters | - |
| `GS1DataBarExpandedStacked` | Linear | String | 74 numeric or 41 alphabetic characters | - |
| `GS1DataBarLimited` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified |
| `GS1DataBarOmnidirectional` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified |
| `GS1DataBarStacked` | Linear | Numbers Only | Up to 14 digits | Last digit must be checksum and will be verified |
| `GTIN12` | Linear | Numbers Only | Expects 11 digits; 12th optional | - |
| `GTIN13` | Linear | Numbers Only | Expects 12 digits; 13th optional | - |
| `GTIN14` | Linear | Numbers Only | Expects 13 digits; 14th optional | - |
| `GTIN8` | Linear | Numbers Only | Expects 7 digits; 8th optional | - |
| `IntelligentMail` | Linear | Numbers Only | Up to 31 digits | Tracking: 20 digits; Rounding: 0,5,9,11 digits; Spaces/dots allowed |
| `Interleaved2of5` | Linear | Numbers Only | - | EVEN if no checksum; ODD if checksum added |
| `ITF14` | Linear | Numbers Only | Expects 13 digits; 14th optional | Will be verified |
| `MaxiCode` | 2D | All ASCII characters | 93 | - |
| `MICR` | - | - | - | - |
| `MicroPDF` | 2D | String | 1850 characters or 2710 digits | - |
| `MSI` | Linear | Numbers Only | Variable | - |
| `PatchCode` | - | - | - | - |
| `PDF417` | Linear | All 256 ASCII characters and 8-bit binary data | Variable | - |
| `Pharmacode` | Linear | Decimal Numbers | 1 to 131070 | - |
| `PostNet` | Postal | Numbers Only | 5 + check digit +4 optional +6 optional | - |
| `PZN` | Linear | Numbers Only | Exactly 6 or 7 digits | - |
| `QRCode` | 2D | All ASCII characters | Variable | - |
| `RoyalMail` | Postal | Digits and characters from A to Z | - | - |
| `RoyalMailKIX` | Postal | All numeric digits (0-9), uppercase letters (A-Z) | Variable | - |
| `Trioptic` | - | - | - | - |
| `UPCA` | Linear | Numbers Only | 11 + check digit +2 optional +5 optional | - |
| `UPCE` | Linear | Numbers Only | 8 digits total | Encodes 6 digits + number system + check digit |
| `UPU` | Postal | - | - | - |
# Barcodes Reader
Source: https://developer.pdf.co/api/barcode/read
Read barcodes from images and **PDF**. Can read all popular barcode types from QR Code and Code 128, EAN to Datamatrix, PDF417, GS1 and many other barcodes.
**Try it live:** [Barcodes Reader → API Tester](/api-tester/barcode/read) — send a real request from your browser.
## `POST /v1/barcode/read/from/url`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `type` | string | *Yes* | QRCode | Set the barcode type to be used. See available barcode types in the [Supported Barcode Types](/api/barcode/overview#supported-barcode-types) |
| `types` | string | *No* | - | Detects checkboxes, radiobuttons, vertical and horizontal lines, and general segments (all content types) on scanned documents using the barcode reader engine. Comma-separated list of object types to decode, must be in a string format. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `RenderingResolution` | integer | *No* | 120 | Set the rendering resolution for the barcode reader engine. The default resolution is 120 DPI. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `barcodes` | array\[object] | List of barcodes found in the document |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
### types
Detects checkboxes, radiobuttons, vertical and horizontal lines, and general segments (all content types) on scanned documents using the barcode reader engine.
Comma-separated list of object types to decode, must be in a string format.
## Visual Element Detection Modes
* Checkbox: Locates check boxes.
* Segment: Locates and selects objects on a page (general selection).
* UnderlinedField: Detects fillable fields (typically, underlined spaces, i.e. fields to fill in a form).
* Rectangle: Detects rectangles, including checkboxes. Also returns the value as 1 if a checkmark or a filled rectangle was detected.
* Oval: Detects rounded or oval marks (typically, a radiobutton). Returns value of 1 if filled out radiobutton was detected.
* HorizontalLine: Detects horizontal lines.
* VerticalLine: Detects vertical lines.
## Normal Example
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf",
"types": "Checkbox,UnderlinedField",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"barcodes": [
{
"Value": "abcdef123456",
"RawData": "",
"Type": 14,
"Rect": "{X=448,Y=23,Width=106,Height=112}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 1,
"Metadata": "",
"TypeName": "QRCode"
},
{
"Value": "test123",
"RawData": "",
"Type": 2,
"Rect": "{X=111,Y=60,Width=255,Height=37}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 0.90625155,
"Metadata": "",
"TypeName": "Code128"
},
{
"Value": "123456",
"RawData": "",
"Type": 4,
"Rect": "{X=111,Y=129,Width=306,Height=37}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 0.7710818,
"Metadata": "",
"TypeName": "Code39"
},
{
"Value": "0112345678901231",
"RawData": "",
"Type": 2,
"Rect": "{X=111,Y=198,Width=305,Height=37}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 0.9156459,
"Metadata": "",
"TypeName": "Code128"
},
{
"Value": "12345670",
"RawData": [
1,
2,
3,
4,
5,
6,
7,
0
],
"Type": 5,
"Rect": "{X=111,Y=267,Width=182,Height=0}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 1,
"Metadata": "",
"TypeName": "I2of5"
},
{
"Value": "1234567890128",
"RawData": "",
"Type": 6,
"Rect": "{X=102,Y=336,Width=71,Height=72}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"Confidence": 0.895925164,
"Metadata": "",
"TypeName": "EAN13"
}
],
"pageCount": 1,
"error": false,
"status": 200,
"remainingCredits": 99826192,
"credits": 35
}
```
#### `Example` CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/barcode/read/from/url' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"types": "QRCode,Code128,Code39,Interleaved2of5,EAN13",
"pages": "0",
"async": false
}'
```
## Optical Marks Reader
Our barcode reader engine can also find the following marks and objects on scanned documents:
* Checkboxes
* Radioboxes
* Vertical and horizontal lines
* General segments (basically, all content types on the page).
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf",
"types": "QRCode,Code128,Code39,Interleaved2of5,EAN13",
"pages": "0",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"barcodes": [
{
"Value": "box",
"RawData": "",
"Type": 53,
"Rect": "{X=298,Y=437,Width=132,Height=6}",
"Page": 0,
"File": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf",
"Confidence": 1,
"Metadata": "",
"TypeName": "UnderlinedField"
}
],
"pageCount": 1,
"error": false,
"status": 200,
"duration": 860,
"remainingCredits": 98725528,
"credits": 35
}
```
#### `Example` CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/barcode/read/from/url' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf",
"types": "Checkbox,UnderlinedField",
"async": false
}'
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```javascript theme={null}
var https = require("https");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source file to search barcodes in.
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf";
// Comma-separated list of barcode types to search.
// See valid barcode types in the documentation https://developer.pdf.co
const BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// Prepare request to `Barcode Reader` API endpoint
var queryPath = `/v1/barcode/read/from/url`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
types: BarcodeTypes,
pages: Pages,
url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Display found barcodes in console
data.barcodes.forEach((element) => {
console.log("Found barcode:");
console.log(" Type: " + element.TypeName);
console.log(" Value: " + element.Value);
console.log(" Document Page Index: " + element.Page);
console.log(" Rectangle: " + element.Rect);
console.log(" Confidence: " + element.Confidence);
console.log("");
}, this);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.error(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source file to search barcodes in.
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf"
# Comma-separated list of barcode types to search.
# See valid barcode types in the documentation https://developer.pdf.co
BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
def main(args=None):
readBarcodes(SourceFileURL)
def readBarcodes(uploadedFileUrl):
"""Get Barcode Information using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {}
parameters["types"] = BarcodeTypes
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'Barcode Reader' API request
url = "{}/barcode/read/from/url".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Display information
for barcode in json["barcodes"]:
print("Found barcode:")
print(f" Type: {barcode['TypeName']}")
print(f" Value: {barcode['Value']}")
print(f" Document Page Index: {barcode['Page']}")
print(f" Rectangle: {barcode['Rect']}")
print(f" Confidence: {barcode['Confidence']}")
print("")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source file (image or PDF) to search barcodes in.
const string SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf";
// Comma-separated list of barcode types to search.
// See valid barcode types in the documentation https://developer.pdf.co
const string BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare requests params as JSON
// See documentation: https://developer.pdf.co
Dictionary parameters = new Dictionary();
parameters.Add("url", SourceFileURL);
parameters.Add("type", BarcodeTypes);
parameters.Add("pages", Pages);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// URL of "Barcode Reader" endpoint
string url = "https://api.pdf.co/v1/barcode/read/from/url";
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Display found barcodes in console
foreach (JToken token in json["barcodes"])
{
Console.WriteLine("Found barcode:");
Console.WriteLine(" Type: " + token["TypeName"]);
Console.WriteLine(" Value: " + token["Value"]);
Console.WriteLine(" Document Page Index: " + token["Page"]);
Console.WriteLine(" Rectangle: " + token["Rect"]);
Console.WriteLine(" Confidence: " + token["Confidence"]);
Console.WriteLine();
}
}
else
{
// Display service reported error
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
// Display request error
Console.WriteLine(e.ToString());
}
finally
{
webClient.Dispose();
}
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source file to search barcodes in.
final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/barcode-reader/sample.pdf";
// Comma-separated list of barcode types to search.
// See valid barcode types in the documentation https://developer.pdf.co
final static String BarcodeTypes = "Code128,Code39,Interleaved2of5,EAN13";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Barcode Reader` API call
String query = "https://api.pdf.co/v1/barcode/read/from/url";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"types\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
BarcodeTypes,
Pages,
SourceFileURL);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Display found barcodes in console
for (JsonElement element : json.get("barcodes").getAsJsonArray())
{
JsonObject barcode = (JsonObject) element;
System.out.println("Found barcode:");
System.out.println(" Type: " + barcode.get("TypeName").getAsString());
System.out.println(" Value: " + barcode.get("Value").getAsString());
System.out.println(" Document Page Index: " + barcode.get("Page").getAsString());
System.out.println(" Rectangle: " + barcode.get("Rect").getAsString());
System.out.println(" Confidence: " + barcode.get("Confidence").getAsString());
System.out.println();
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
}
```
```php theme={null}
Cloud API asynchronous "Barcode Reader" job example (allows to avoid timeout errors).
";
if (curl_errno($curl) == 0)
{
$status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE);
if ($status_code == 200)
{
$json = json_decode($result, true);
if (!isset($json["error"]) || $json["error"] == false)
{
// URL of generated JSON file that will available after the job completion
$resultFileUrl = $json["url"];
// Asynchronous job ID
$jobId = $json["jobId"];
// Check the job status in a loop
do
{
$status = CheckJobStatus($jobId, $apiKey); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
echo "
" . date(DATE_RFC2822) . ": " . $status . "
";
if ($status == "success")
{
// Display link to JSON file with information about decoded barcodes
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Excel to CSV
Source: https://developer.pdf.co/api/convert-from-excel/csv
Converts a `xls`/`xlsx` file to `csv`.
**Try it live:** [Excel to CSV → API Tester](/api-tester/convert-from-excel/csv) — send a real request from your browser.
## `POST /v1/xls/convert/to/csv`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/xls/convert/to/csv' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false
}'
```
```python theme={null}
import requests
import json
# Your API endpoint URL.
url = "https://api.pdf.co/v1/xls/convert/to/csv"
# Your API Key.
api_key = "Your API Key"
# The URL of the Excel file you want to convert.
input_file_url = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls"
headers = {
"x-api-key": api_key,
"Content-Type": "application/json"
}
data = {
"url": input_file_url,
"async": False
}
response = requests.post(url, headers=headers, json=data)
if response.status_code == 200:
# The request was successful.
# Parse the json response.
data = response.json()
# Extract the CSV file URL from the response.
csv_url = data.get('url', '')
print("CSV file is available at: ", csv_url)
else:
# There was an error with the request.
print("Error: ", response.status_code)
```
# Excel to HTML
Source: https://developer.pdf.co/api/convert-from-excel/html
Converts a `xls`/`xlsx`/`csv` file to `html`.
**Try it live:** [Excel to HTML → API Tester](/api-tester/convert-from-excel/html) — send a real request from your browser.
## `POST /v1/xls/convert/to/html`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
# Excel to JSON
Source: https://developer.pdf.co/api/convert-from-excel/json
Converts a `xls`/`xlsx`/`csv` file to `json`.
**Try it live:** [Excel to JSON → API Tester](/api-tester/convert-from-excel/json) — send a real request from your browser.
## `POST /v1/xls/convert/to/json`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
# Excel to PDF
Source: https://developer.pdf.co/api/convert-from-excel/pdf
Converts a `xls`/`xlsx`/`csv` file to `pdf`.
**Try it live:** [Excel to PDF → API Tester](/api-tester/convert-from-excel/pdf) — send a real request from your browser.
## `POST /v1/xls/convert/to/pdf`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination CSV file name
const DestinationFile = "./result.csv";
// Prepare request to `PDF To CSV` API endpoint
var queryPath = `/v1/xls/convert/to/csv`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download CSV file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated CSV file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
import json
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source xls file.
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls"
def main(args = None):
convertXlsToXml(SourceFileURL)
def convertXlsToXml(sourceFileUrl):
"""Convert Xls/Xlsx to Xml using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {
"url": sourceFileUrl,
"async": False
}
# Prepare URL for 'Xls to Xml' API request
url = "{}/xls/convert/to/xml".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, headers={ "x-api-key": API_KEY, "Content-Type": "application/json" }, data=json.dumps(parameters))
if (response.status_code == 200):
json_res = response.json()
if json_res["error"] == False:
# Get URL of result file
resultFileUrl = json_res["url"]
# Output URL of converted xml file
print(f"Result file url: {resultFileUrl}")
else:
# Show service reported error
print(json_res["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
# Excel to Text
Source: https://developer.pdf.co/api/convert-from-excel/text
Converts a `xls`/`xlsx`/`csv` file to `text`.
**Try it live:** [Excel to Text → API Tester](/api-tester/convert-from-excel/text) — send a real request from your browser.
## `POST /v1/xls/convert/to/txt`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
# Excel to XML
Source: https://developer.pdf.co/api/convert-from-excel/xml
Converts a `xls`/`xlsx`/`csv` file to `xml`.
**Try it live:** [Excel to XML → API Tester](/api-tester/convert-from-excel/xml) — send a real request from your browser.
## `POST /v1/xls/convert/to/xml`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `worksheetIndex` | string | *No* | - | Set the index of the worksheet to be used. The first worksheet has index 1. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"async": false,
"name": "Output"
}
```
# Credits per API Function
Source: https://developer.pdf.co/api/credits-per-api-function
Estimate PDF.co API credit usage by endpoint. See how many credits each API function uses and whether the charge is per page or per API call.
Every PDF.co API endpoint consumes credits when you call it. This page lists the credit cost of each endpoint so you can estimate how many credits a workflow will use before you run it. Endpoints are charged in one of two ways, **per processed page** or **per API call**, and the tables below are grouped accordingly.
For a quick estimate, use the interactive **Credit Calculator** on the [Subscriptions](https://app.pdf.co/subscriptions) page. It lets you pick an endpoint and page count and shows the credit cost instantly.
## How to estimate credits
| Usage type | Formula |
| ------------ | ---------------------------------------------------- |
| Per page | `total credits = endpoint credits × pages processed` |
| Per API call | `total credits = endpoint credits × API calls` |
For example, converting 5,000 one-page PDFs to JPG with `/v1/pdf/convert/to/jpg` uses `5,000 × 12 = 60,000 credits`.
If each document has multiple pages, multiply by the total number of pages. For example, 5,000 documents with 3 pages each is 15,000 pages, so PDF to JPG would use `15,000 × 12 = 180,000 credits`.
## Endpoints charged per page
| Endpoint | API path | Credits (per page) |
| ------------------------------------------------ | -------------------------------------- | -----------------: |
| Parse invoices with AI | `/v1/ai-invoice-parser` | 100 |
| Extract data with Document Parser template | `/v1/pdf/documentparser` | 42 |
| Add text, images, and fill form fields | `/v1/pdf/edit/add` | 21 |
| Merge PDFs into one | `/v1/pdf/merge` | 2 |
| Merge images, documents, and PDFs into a new PDF | `/v1/pdf/merge2` | 35 |
| Split PDF by page numbers | `/v1/pdf/split` | 2 |
| Split PDF by text search | `/v1/pdf/split2` | 35 |
| Convert XLS or XLSX to PDF | `/v1/xls/convert/to/pdf` | 21 |
| Convert CSV to PDF | `/v1/pdf/convert/from/csv` | 21 |
| Convert DOC, DOCX, RTF, TXT, or XPS to PDF | `/v1/pdf/convert/from/doc` | 21 |
| Convert HTML to PDF | `/v1/pdf/convert/from/html` | 9 |
| Convert images to PDF | `/v1/pdf/convert/from/image` | 9 |
| Convert URL to PDF | `/v1/pdf/convert/from/url` | 9 |
| Convert EML or MSG to PDF | `/v1/pdf/convert/from/email` | 56 |
| Convert PDF to CSV (AI-powered) | `/v1/pdf/convert/to/csv` | 28 |
| Convert PDF to HTML | `/v1/pdf/convert/to/html` | 21 |
| Convert PDF to JSON (legacy) | `/v1/pdf/convert/to/json` | 28 |
| Convert PDF to JSON (AI-powered) | `/v1/pdf/convert/to/json2` | 28 |
| Convert PDF to JSON with metadata (AI-powered) | `/v1/pdf/convert/to/json-meta` | 42 |
| Convert PDF to text (AI-powered) | `/v1/pdf/convert/to/text` | 21 |
| Convert PDF to text (simple, no AI) | `/v1/pdf/convert/to/text-simple` | 4 |
| Convert PDF to XLS (AI-powered) | `/v1/pdf/convert/to/xls` | 35 |
| Convert PDF to XLSX (AI-powered) | `/v1/pdf/convert/to/xlsx` | 28 |
| Convert PDF to XML (AI-powered) | `/v1/pdf/convert/to/xml` | 35 |
| Render PDF to JPG | `/v1/pdf/convert/to/jpg` | 12 |
| Render PDF to PNG | `/v1/pdf/convert/to/png` | 15 |
| Render PDF to WebP | `/v1/pdf/convert/to/webp` | 18 |
| Render PDF to TIFF | `/v1/pdf/convert/to/tiff` | 28 |
| Convert XLS or XLSX to CSV | `/v1/xls/convert/to/csv` | 9 |
| Convert XLS or XLSX to HTML | `/v1/xls/convert/to/html` | 9 |
| Convert XLS or XLSX to JSON | `/v1/xls/convert/to/json` | 15 |
| Convert XLS or XLSX to TXT | `/v1/xls/convert/to/txt` | 9 |
| Convert XLS or XLSX to XML | `/v1/xls/convert/to/xml` | 15 |
| Rotate pages | `/v1/pdf/edit/rotate` | 7 |
| Detect and fix page rotation | `/v1/pdf/edit/rotate/auto` | 28 |
| Remove pages from PDF | `/v1/pdf/edit/delete-pages` | 5 |
| Replace text in PDF | `/v1/pdf/edit/replace-text` | 21 |
| Replace text with an image in PDF | `/v1/pdf/edit/replace-text-with-image` | 77 |
| Delete text in PDF | `/v1/pdf/edit/delete-text` | 21 |
| Read barcodes from URL or file | `/v1/barcode/read/from/url` | 35 |
| Extract attachments from MSG or EML | `/v1/email/extract-attachments` | 35 |
| Decode email from MSG or EML | `/v1/email/decode` | 35 |
| Send email with attachments | `/v1/email/send` | 21 |
| Add security protection to PDF | `/v1/pdf/security/add` | 3 |
| Remove protection from PDF | `/v1/pdf/security/remove` | 3 |
| Extract PDF attachments | `/v1/pdf/attachments/extract` | 8 |
| Classify document based on rules | `/v1/pdf/classifier` | 42 |
| Compress PDF to reduce file size | `/v2/pdf/compress` | 35 |
| Read PDF file information | `/v1/pdf/info` | 7 |
| Find text inside PDFs and images | `/v1/pdf/find` | 35 |
| Return JSON with table information | `/v1/pdf/find/table` | 21 |
| Convert scanned PDF to text-searchable PDF | `/v1/pdf/makesearchable` | 35 |
| Convert PDF to scanned (unsearchable) PDF | `/v1/pdf/makeunsearchable` | 35 |
## Endpoints charged per API call
| Endpoint | API path | Credits (per API call) |
| --------------------------------------------- | -------------------------------------- | ---------------------: |
| Generate barcode image | `/v1/barcode/generate` | 7 |
| Generate file upload URL | `/v1/file/upload/get-presigned-url` | 7 |
| Upload a small local file as a temporary file | `/v1/file/upload` | 11 |
| Upload file from URL | `/v1/file/upload/url` | 11 |
| Upload file from Base64 | `/v1/file/upload/base64` | 21 |
| Get HTML templates | `/v1/templates/html` | 2 |
| Get Document Parser templates | `/v1/pdf/documentparser/templates` | 2 |
| Get Document Parser template by ID | `/v1/pdf/documentparser/templates/:id` | 2 |
| Read PDF form fields | `/v1/pdf/info/fields` | 8 |
| Check background job status | `/v1/job/check` | 2 |
## Notes
* [Estimated credits](/api/async-and-sync-mode) are calculated for sync mode and could differ in async mode.
* Some workflows make more than one API call. For example, uploading a file separately before processing it consumes upload endpoint credits in addition to the processing endpoint credits.
* Async jobs may require status checks with `/v1/job/check`, which is listed separately in the per API call table.
* To check the remaining credits in an account, use [`/v1/account/credit/balance`](/api/account-balance-info).
* You can also estimate costs with the [Credit Calculator](https://app.pdf.co/subscriptions) on the Subscriptions page.
# Document Classifier
Source: https://developer.pdf.co/api/document-classifier
Detect the class of an incoming PDF, JPG, or PNG document using keyword rules or built-in AI, so you can route it to the right processing template.
**Try it live:** [Document Classifier → API Tester](/api-tester/document-classifier) — send a real request from your browser.
## `POST /v1/pdf/classifier`
Document Classifier can automatically find class of input PDF, JPG, PNG document by analyzing its content using the built-in AI or custom defined classification rules.
The best way to **develop**, **test** and **maintain** classification rules is to use `Classifier Tester Tool` from PDF.co [Document Classifier UI](https://app.pdf.co/document-classifier) . Use this tool to quickly edit and test rules on single PDFs and on folders.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. |
| `rulescsv` | string | *No* | - | Define custom classification rules in CSV format. See the [rulescsv](#rulescsv). |
| `rulescsvurl` | string | *No* | - | URL to the CSV file with classification rules. For the format, see the description above `rulescsv` parameter |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `RenderTextObjects` | boolean | *No* | `true` | Render text objects or not |
| `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not |
| `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not |
| `TIFFCompression` | string | *No* | `LZW` | TIFF compression algorithm. The options are: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE` |
| `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters.AddGammaCorrection()` | array\[string (float format)] | *No* | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. |
| `OCRImagePreprocessingFilters.AddGrayscale()` | boolean | *No* | `false` | Set to true to preprocessing filter that converts a colored document/image to grayscale before performing OCR |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### `rulescsv`
Rules are in CSV format where each row contains: `class name`, `logic` (`AND` or `OR` (default)), and keywords separated by a comma. Each row is separated by the `\n` symbol. You can use regular expressions for keywords with this syntax: `/keyword or regexp/i` where `i` is the case-insensitive flag. Please note that all `\` symbols should add the prefix `\` because of JSON format, so `\d` becomes `\\d` and so on.
> **Custom Rules Example 1** for `rulescsv`.
>
> ```
> Amazon AWS, OR, Amazon Web Services Invoice, Amazon CloudFront\nDigital Ocean, OR,DigitalOcean, DOInvoice\nACME,OR, ACME Inc.,1540 Long Street
> ```
> **Custom Rules Example 2**.
>
> ```
> Medical Report,AND,/Instructing Party|Medical Report|Date Of Injury|Med Agency Ref/i\r\nInjured Claimant,OR, Injured Claimant, Injured Patient ID
> ```
## Document Classifier Usage Guide
This Document Classifier checks content of input PDF, JPG, PNG, or TIFF. It uses AI to automatically determine the class of the document (e.g., `finance`, `invoice`) and returns the result to the user. Custom-defined classification rules can also be used.
Use this Document Classifier to quickly build a workflow for sorting input documents and PDF files.
### How to Create and Test Custom Classification Rules
Classification rules are stored in CSV format, one line per class, with the following format:
```
className, logicType, keyword1, keyword2, keyword3 ...
```
Where:
* `className` – The name of the class. It will be returned if rules from this class match the document.
* `logicType` – (Optional) Logic to use for keywords. Can be `OR` (default) or `AND`. `OR` means the class is identified if one or more keywords match. `AND` means **all** keywords must match. If not specified, `OR` is assumed.
* `keyword1`, `keyword2`, `keyword3` – Keywords or phrases to check. Can include regular expressions, e.g., `/\d+/` or `/Medical Report|Med Report/i`.
### Sample Rules
```
Invoice,OR,Invoice Number,Invoice #,Invoice No,Tax Invoice,,
Purchase Order,OR,PO Number,Order Number,Order No,,,
Bill,OR,Bill Date,Billing Period,Bill Number,,,
Bank Statement,OR,/Account Statement/i,/Statement of Account/i,Business Checking,Accounts Payable,/Statement No/i,
Income Statement,OR,/Income Statement/i,,,,,
Has US Number,OR,"/\b-?(\d+,?)+(\.\d\d)\b/",,,,,
Medical Report,AND,/Medical Report|Med Report/i
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `body` | object | Response body. |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes).. For more information, see [Response Codes](/api/response-codes). |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"async": false,
"inline": "true",
"password": "",
"profiles": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"classes": [
{
"class": "invoice"
},
{
"class": "finance"
},
{
"class": "documents"
}
]
},
"pageCount": 1,
"error": false,
"status": 200,
"credits": 42,
"duration": 353,
"remainingCredits": 98019328
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/classifier' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"async": false,
"inline": "true",
"password": "",
"profiles": ""
} '
```
```javascript theme={null}
var request = require('request');
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/pdf/classifier',
'headers': {
'Content-Type': 'application/json',
'x-api-key': 'YOUR_PDFCO_API_KEY'
},
body: JSON.stringify({
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"async": false,
"encrypt": "false",
"inline": "true",
"password": "",
"profiles": ""
})
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
```csharp theme={null}
using System;
using RestSharp;
namespace HelloWorldApplication {
class HelloWorld {
static void Main(string[] args) {
var client = new RestClient("https://api.pdf.co/v1/pdf/classifier");
client.Timeout = -1;
var request = new RestRequest(Method.POST);
request.AddHeader("Content-Type", "application/json");
request.AddHeader("x-api-key", "YOUR_PDFCO_API_KEY");
var body = @"{" + "\n" +
@" ""url"": ""https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf""," + "\n" +
@" ""async"": false," + "\n" +
@" ""encrypt"": ""false""," + "\n" +
@" ""inline"": ""true""," + "\n" +
@" ""password"": """"," + "\n" +
@" ""profiles"": """"" + "\n" +
@"} ";
request.AddParameter("application/json", body, ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
Console.WriteLine(response.Content);
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("application/json");
RequestBody body = RequestBody.create(mediaType, "{\n \"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf\",\n \"async\": false,\n \"encrypt\": \"false\",\n \"inline\": \"true\",\n \"password\": \"\",\n \"profiles\": \"\"\n} ");
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/classifier")
.method("POST", body)
.addHeader("Content-Type", "application/json")
.addHeader("x-api-key", "YOUR_PDFCO_API_KEY")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
'https://api.pdf.co/v1/pdf/classifier',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => '',
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_POSTFIELDS =>'{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"async": false,
"encrypt": "false",
"inline": "true",
"password": "",
"profiles": ""
} ',
CURLOPT_HTTPHEADER => array(
'Content-Type: application/json',
'x-api-key: YOUR_PDFCO_API_KEY'
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
?>
```
# Document Parser Overview
Source: https://developer.pdf.co/api/documentparser/overview
Parse PDFs and scanned documents to extract fields, tables, values, and barcodes from invoices, statements, orders, and similar forms.
**Try it live:** [Parse Document → API Tester](/api-tester/documentparser) — send a real request from your browser.
## Built-in document parser templates
`General Invoice Template` can parse invoices (English only) to invoice id, invoice date, extract total, tax, and line items. Set the `templateId` parameter to `1` to use this template.
## How to classify incoming documents before parsing them?
Use the [/pdf/classifier](/api/document-classifier) endpoint (see below) to automatically sort/detect the class of the document based on AI or on custom keywords-based rules.
For example, you can easily define rules to find which vendor provided the document to find which template to apply accordingly. See [Document Classifier](https://developer.pdf.co/api/document-classifier) for more details.
## Additional Information and Tools
* [Document Parser Template Editor](https://app.pdf.co/document-parser/templates)
* [PDF.co Document Parser: Template Creation Guide](https://developer.pdf.co/knowledgebase/document-parser-guide)
# Parse Document
Source: https://developer.pdf.co/api/documentparser/parser
Extract fields, tables, and values from PDFs and images using Document Parser templates with searchable form fields and multi-page support.
## `POST /v1/pdf/documentparser`
Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `templateId` | integer | *No* | - | Set ID of document parser template to be used. View and manage your templates at [Document Parser](https://app.pdf.co/document-parser/templates) |
| `template` | string | *No* | - | The raw format of the document parser template to be used directly. see [Template](/api/documentparser/parser) |
| `password` | string | *No* | - | Password for the PDF file. |
| `inline` | boolean | *No* | `false` | Set to true to include the results directly in the response, in addition to providing a URL to the generated output file. Applies only when `async` mode is enabled. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `outputFormat` | string | *No* | `JSON` | The format of the output file. The output format can be `JSON`, `CSV`, or `XML`. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
| `body` | object | *No* |
| `objects` | array\[object] | |
| `elapsed` | float | Processing time in seconds |
| `templateName` | string | Name of the parsing template used |
| `templateVersion` | string | Version of the parsing template |
| `timestamp` | string | Timestamp when the parsing occurred |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"outputFormat": "JSON",
"templateId": "1",
"async": false,
"inline": "true",
"password": "",
"profiles": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"objects": [
{
"name": "companyName",
"objectType": "field",
"value": "Amazon Web Services, Inc",
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "companyName2",
"objectType": "field",
"value": "Amazon Web Services, Inc",
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "invoiceId",
"objectType": "field",
"value": "123456789",
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "dateIssued",
"objectType": "field",
"value": "2018-04-03T00:00:00",
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "dateDue",
"objectType": "field",
"value": "2018-04-03T00:00:00",
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "bankAccount",
"objectType": "field",
"value": "123456789012",
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "total",
"objectType": "field",
"value": 6.58,
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"name": "subTotal",
"objectType": "field",
"value": ""
},
{
"name": "tax",
"objectType": "field",
"value": 1.01,
"pageIndex": 0,
"rectangle": [
0,
0,
0,
0
]
},
{
"objectType": "table",
"name": "table",
"rows": []
}
],
"templateName": "Generic Invoice [en]",
"templateVersion": "4",
"timestamp": "2020-08-21T19:23:31"
},
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample-invoice.json",
"remainingCredits": 60803
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/documentparser' \
--header 'Content-Type: application/json' \
--header 'x-api-key: {{x-api-key}}' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf",
"outputFormat": "JSON",
"templateId": "1",
"async": false,
"inline": "true",
"password": "",
"profiles": ""
}'
```
```javascript theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
// Source PDF file
final Path SourceFile = Paths.get(".\\MultiPageTable.pdf");
// PDF document password. Leave empty for unprotected documents.
final String Password = "";
// Destination JSON file name
final Path DestinationFile = Paths.get(".\\result.json");
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8);
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. PARSE UPLOADED PDF DOCUMENT
ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String uploadedFileUrl, String templateText) throws IOException
{
// Prepare POST request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("url", new JsonPrimitive(uploadedFileUrl));
jsonBody.add("template", new JsonPrimitive(templateText));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request to `Document Parser` API
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/documentparser")
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated JSON file
String resultFileUrl = json.get("url").getAsString();
// Download JSON file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\MultiPageTable.pdf"
# Destination JSON file name
DestinationFile = ".\\result.json"
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
# to create templates.
# Read template from file:
file_read = open(".\\MultiPageTable-template1.yml", mode='r', encoding="utf-8",errors="ignore")
Template = file_read.read()
file_read.close()
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
PerformDocumentParser(uploadedFileUrl, Template, DestinationFile)
def PerformDocumentParser(uploadedFileUrl, template, destinationFile):
# Content
data = {
'url': uploadedFileUrl,
'template': template
}
# Prepare URL for 'Document Parser' API request
url = "{}/pdf/documentparser".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data= data, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file,
headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"})
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\MultiPageTable.pdf";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination TXT file name
const string DestinationFile = @".\result.json";
static void Main(string[] args)
{
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
String templateText = File.ReadAllText(@".\MultiPageTable-template1.yml");
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. PARSE UPLOADED PDF DOCUMENT
// URL of `Document Parser` API call
string url = "https://api.pdf.co/v1/pdf/documentparser";
Dictionary requestBody = new Dictionary();
requestBody.Add("template", templateText);
requestBody.Add("name", Path.GetFileName(DestinationFile));
requestBody.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(requestBody);
// Execute request
response = webClient.UploadString(url, "POST", jsonPayload);
// Parse response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated JSON file
string resultFileUrl = json["url"].ToString();
// Download JSON file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
// Source PDF file
final Path SourceFile = Paths.get(".\\MultiPageTable.pdf");
// PDF document password. Leave empty for unprotected documents.
final String Password = "";
// Destination JSON file name
final Path DestinationFile = Paths.get(".\\result.json");
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8);
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. PARSE UPLOADED PDF DOCUMENT
ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String uploadedFileUrl, String templateText) throws IOException
{
// Prepare POST request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("url", new JsonPrimitive(uploadedFileUrl));
jsonBody.add("template", new JsonPrimitive(templateText));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request to `Document Parser` API
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/documentparser")
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated JSON file
String resultFileUrl = json.get("url").getAsString();
// Download JSON file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Document Parse Results
Status code: " . $status_code . "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# List All Templates
Source: https://developer.pdf.co/api/documentparser/templates
Returns all **Document Parser** data extraction templates available to the current user.
**Try it live:** [List All Templates → API Tester](/api-tester/documentparser/templates) — send a real request from your browser.
## `GET /v1/pdf/documentparser/templates`
Use the PDF.co dashbaord to manage your [Document Parser Templates](https://app.pdf.co/document-parser/templates).
Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | -------------- | ------------------------------------------ |
| `templates` | array\[object] | |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `credits` | integer | Number of credits consumed by the request |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"templates": [
{
"id": 40,
"type": "user",
"title": "Untitled",
"description": "Untitled"
},
{
"id": 1,
"type": "system",
"title": "Invoice Parser",
"description": "Parses invoices and extracts invoice number, company name, due date, amount, tax"
}
],
"remainingCredits": 94229
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/pdf/documentparser/templates' \
--header 'Content-Type: application/json' \
--header 'x-api-key: {{x-api-key}}'
```
# Retrieve Template by ID
Source: https://developer.pdf.co/api/documentparser/templates-id
Returns detailed information for document parser template by template's id.
**Try it live:** [Retrieve Template by ID → API Tester](/api-tester/documentparser/templates-id) — send a real request from your browser.
## `GET /v1/pdf/documentparser/templates/:id`
Use the PDF.co dashbaord to manage your [Document Parser Templates](https://app.pdf.co/document-parser/templates).
Please refer to the [Document Parser Template Editor](https://app.pdf.co/document-parser/templates) and [PDF.co Document Parser: Template Creation Guide](/knowledgebase/document-parser-guide#document-parser-template-objects-guide) for more information.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ----------------- | -------------------------------------------------------------------- |
| `id` | integer | Unique identifier for the template |
| `type` | string | Source of the template. The available sources are: `user`, `system`. |
| `title` | string | Title of the template |
| `description` | string | Description of what the template does |
| `created_at` | String (ISO 8601) | Timestamp indicating when the template was initially created |
| `updated_at` | String (ISO 8601) | Timestamp indicating the last time the template was modified |
| `body` | string | Template content |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/pdf/documentparser/templates/1' \
--header 'Content-Type: application/json' \
--header 'x-api-key: {{x-api-key}}' \
--data-raw ''
```
```javascript theme={null}
/*jshint esversion: 6 */
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./MultiPageTable.pdf";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.json";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. OPTIMIZE UPLOADED PDF FILE
parsePdf(API_KEY, uploadedFileUrl, Password, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function parsePdf(apiKey, uploadedFileUrl, password, destinationFile) {
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
var templateText = fs.readFileSync("./MultiPageTable-template1.yml", "utf-8");
// URL for `Document Parser` API call
var query = `https://api.pdf.co/v1/pdf/documentparser`;
var jsonRequestObject = {
url: uploadedFileUrl,
template: templateText
};
request(
{
url: query,
headers: { "x-api-key": API_KEY },
method: "POST",
json: true,
body: jsonRequestObject
},
function (error, response, body) {
if (error) {
return console.error("Error: ", error);
}
// Parse JSON response
let data = JSON.parse(JSON.stringify(body));
if (data.error == false) {
//Download generated file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated result file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("Error: " + data.message);
}
}
);
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\MultiPageTable.pdf"
# Destination JSON file name
DestinationFile = ".\\result.json"
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
# to create templates.
# Read template from file:
file_read = open(".\\MultiPageTable-template1.yml", mode='r', encoding="utf-8",errors="ignore")
Template = file_read.read()
file_read.close()
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
PerformDocumentParser(uploadedFileUrl, Template, DestinationFile)
def PerformDocumentParser(uploadedFileUrl, template, destinationFile):
# Content
data = {
'url': uploadedFileUrl,
'template': template
}
# Prepare URL for 'Document Parser' API request
url = "{}/pdf/documentparser".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data= data, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file,
headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"})
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\MultiPageTable.pdf";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination TXT file name
const string DestinationFile = @".\result.json";
static void Main(string[] args)
{
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
String templateText = File.ReadAllText(@".\MultiPageTable-template1.yml");
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. PARSE UPLOADED PDF DOCUMENT
// URL of `Document Parser` API call
string url = "https://api.pdf.co/v1/pdf/documentparser";
Dictionary requestBody = new Dictionary();
requestBody.Add("template", templateText);
requestBody.Add("name", Path.GetFileName(DestinationFile));
requestBody.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(requestBody);
// Execute request
response = webClient.UploadString(url, "POST", jsonPayload);
// Parse response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated JSON file
string resultFileUrl = json["url"].ToString();
// Download JSON file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
// Source PDF file
final Path SourceFile = Paths.get(".\\MultiPageTable.pdf");
// PDF document password. Leave empty for unprotected documents.
final String Password = "";
// Destination JSON file name
final Path DestinationFile = Paths.get(".\\result.json");
// Template text. Use Document Parser (https://pdf.co/document-parser, https://app.pdf.co/document-parser)
// to create templates.
// Read template from file:
String templateText = new String(Files.readAllBytes(Paths.get(".\\MultiPageTable-template1.yml")), StandardCharsets.UTF_8);
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. PARSE UPLOADED PDF DOCUMENT
ParseDocument(webClient, API_KEY, DestinationFile, Password, uploadedFileUrl, templateText);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void ParseDocument(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String uploadedFileUrl, String templateText) throws IOException
{
// Prepare POST request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("url", new JsonPrimitive(uploadedFileUrl));
jsonBody.add("template", new JsonPrimitive(templateText));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request to `Document Parser` API
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/documentparser")
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated JSON file
String resultFileUrl = json.get("url").getAsString();
// Download JSON file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Document Parse Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Extract Data from Email File
Source: https://developer.pdf.co/api/email/decode
Decode an email message to extract its components.
**Try it live:** [Extract Data from Email File → API Tester](/api-tester/email/decode) — send a real request from your browser.
## `POST /v1/email/decode`
For converting email to PDF please see [PDF from Email](/api/pdf-from-email).
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `responseParameters` | object | *No* | - | - |
| `body` | object | *No* | - | Response body. |
| `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) |
| `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | *No* | - | Name of the output file |
| `credits` | integer | *No* | - | Number of credits consumed by the request |
| `remainingCredits` | integer | *No* | - | Number of credits remaining in the account |
| `duration` | integer | *No* | - | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml",
"inline": true,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"from": "test@example.com",
"fromName": "",
"to": [
{
"address": "test2@example.com",
"name": ""
}
],
"cc": [],
"bcc": [],
"sentAt": null,
"receivedAt": null,
"subject": "Test email with attachments",
"bodyHtml": null,
"bodyText": "Test Email Message with 2 PDF files as attachments\r\n\r\n",
"attachmentCount": 2
},
"error": false,
"status": 200,
"name": "sample.json",
"remainingCredits": 60095
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/email/decode' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml",
"inline": true,
"async": false
}'
```
# Extract Email Attachment
Source: https://developer.pdf.co/api/email/extract-attachments
Extract attachments from an email
**Try it live:** [Extract Email Attachment → API Tester](/api-tester/email/extract-attachments) — send a real request from your browser.
## `POST /v1/email/extract-attachments`
For converting email to PDF please see [PDF from Email](/api/pdf-from-email).
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `responseParameters` | object | *No* | - | - |
| `body` | object | *No* | - | Response body. |
| `pageCount` | integer | *No* | - | Number of pages in the PDF document. |
| `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) |
| `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | *No* | - | Name of the output file |
| `credits` | integer | *No* | - | Number of credits consumed by the request |
| `remainingCredits` | integer | *No* | - | Number of credits remaining in the account |
| `duration` | integer | *No* | - | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml",
"inline": true,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"from": "test@example.com",
"subject": "Test email with attachments",
"bodyHtml": null,
"bodyText": "Test Email Message with 2 PDF files as attachments\r\n\r\n",
"attachments": [
{
"filename": "DigitalOcean.pdf",
"url": "https://pdf-temp-files.s3.amazonaws.com/2943e6bb80e646ec92e839292e95d542/DigitalOcean.pdf"
},
{
"filename": "sample.pdf",
"url": "https://pdf-temp-files.s3.amazonaws.com/e10e37fbb438432a83ece50ccdc719b3/sample.pdf"
}
]
},
"pageCount": 2,
"error": false,
"status": 200,
"name": "sample.json",
"remainingCredits": 60085
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/email/extract-attachments' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml",
"inline": true,
"async": false
}'
```
```javascript theme={null}
var request = require('request');
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/email/send',
'headers': {
'Content-Type': 'application/json',
'x-api-key': 'ADD_YOUR_PDFco_API_KEY'
},
body: JSON.stringify({
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"from": "John Doe ",
"to": "Partner ",
"subject": "Check attached sample pdf",
"bodytext": "Please check the attached pdf",
"bodyHtml": "Please check the attached pdf",
"smtpserver": "smtp.gmail.com",
"smtpport": "587",
"smtpusername": "my@gmail.com",
"smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833",
"async": false
})
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
```python theme={null}
import requests
import json
url = "https://api.pdf.co/v1/email/send"
payload = json.dumps({
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"from": "John Doe ",
"to": "Partner ",
"subject": "Check attached sample pdf",
"bodytext": "Please check the attached pdf",
"bodyHtml": "Please check the attached pdf",
"smtpserver": "smtp.gmail.com",
"smtpport": "587",
"smtpusername": "my@gmail.com",
"smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833",
"async": False
})
headers = {
'Content-Type': 'application/json',
'x-api-key': 'ADD_YOUR_API_KEY'
}
response = requests.request("POST", url, headers=headers, data=payload)
print(response.text)
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.Net;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source PDF file.
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf";
// Email Details
const string From = "John Doe ";
const string To = "Partner ";
const string Subject = "Check attached sample pdf";
const string BodyText = "Please check the attached pdf";
const string BodyHtml = "Please check the attached pdf";
const string SmtpServer = "smtp.gmail.com";
const string SmtpPort = "587";
const string SmtpUserName = "my@gmail.com";
const string SmtpPassword = "app specific password created as https://support.google.com/accounts/answer/185833";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Email Send` API call
string url = "https://api.pdf.co/v1/email/send";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("url", SourceFileUrl);
parameters.Add("from", From);
parameters.Add("to", To);
parameters.Add("subject", Subject);
parameters.Add("bodytext", BodyText);
parameters.Add("bodyHtml", BodyHtml);
parameters.Add("smtpserver", SmtpServer);
parameters.Add("smtpport", SmtpPort);
parameters.Add("smtpusername", SmtpUserName);
parameters.Add("smtppassword", SmtpPassword);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
Console.WriteLine("Email Sent Successfully!");
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("application/json");
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM)
.addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-extractor/sample.eml")
.build();
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/email/extract-attachments")
.method("POST", body)
.addHeader("Content-Type", "application/json")
.addHeader("x-api-key", "{{x-api-key}}")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
'https://api.pdf.co/v1/email/send',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => '',
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_POSTFIELDS =>'{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"from": "John Doe ",
"to": "Partner ",
"subject": "Check attached sample pdf",
"bodytext": "Please check the attached pdf",
"bodyHtml": "Please check the attached pdf",
"smtpserver": "smtp.gmail.com",
"smtpport": "587",
"smtpusername": "my@gmail.com",
"smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833",
"async": false
}',
CURLOPT_HTTPHEADER => array(
'Content-Type: application/json',
'x-api-key: ADD_YOUR_PDFco_KEY_HERE'
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
?>
```
# Send Email with File
Source: https://developer.pdf.co/api/email/send
Send an email. An email can be with or without attachment.
**Try it live:** [Send Email with File → API Tester](/api-tester/email/send) — send a real request from your browser.
## `POST /v1/email/send`
For converting email to PDF please see [PDF from Email](/api/pdf-from-email).
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `from` | string | *Yes* | - | The "From" field with sender name and email |
| `to` | string | *Yes* | - | The "To" field with receiver name and email |
| `subject` | string | *Yes* | - | The subject for the outgoing email. |
| `bodytext` | string | *No* | - | The plain text version of the outgoing email message. |
| `bodyhtml` | string | *No* | - | The HTML version of the outgoing email message. |
| `smtpserver` | string | *Yes* | - | The SMTP server to use for sending the email. |
| `smtpport` | integer | *Yes* | - | The port number of the SMTP server. |
| `smtpusername` | string | *Yes* | - | The username for the SMTP server. |
| `smtppassword` | string | *Yes* | - | The password for the SMTP server. If you use Gmail then you need to generate an [app-specific password](https://support.google.com/accounts/answer/185833) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"from": "John Doe ",
"to": "Partner ",
"subject": "Check attached sample pdf",
"bodytext": "Please check the attached pdf",
"bodyHtml": "Please check the attached pdf",
"smtpserver": "smtp.gmail.com",
"smtpport": "587",
"smtpusername": "my@gmail.com",
"smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"error": false,
"status": 200,
"remainingCredits": 60095
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/email/send' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"from": "John Doe ",
"to": "Partner ",
"subject": "Check attached sample pdf",
"bodytext": "Please check the attached pdf",
"bodyHtml": "Please check the attached pdf",
"smtpserver": "smtp.gmail.com",
"smtpport": "587",
"smtpusername": "my@gmail.com",
"smtppassword": "app specific password created as https://support.google.com/accounts/answer/185833",
"async": false
}'
```
# File Download
Source: https://developer.pdf.co/api/file-download
Download files from PDF.co Files Storage using a unique filetoken; access is restricted to the account that owns the file.
## `GET /v1/file/download/{filetoken}`
**Endpoint URL Format:**
```
https://api.pdf.co/v1/file/download/{filetoken}
```
Replace `{filetoken}` with the actual filetoken identifier. For example:
```
https://api.pdf.co/v1/file/download/a1d30e75adf5eaa.................
```
**Key Features:**
* **Exclusive Access:** The endpoint strictly controls access, allowing only the account owner with the correct filetoken to retrieve the associated file.
* **Secure File Retrieval:** Files are securely stored and can only be accessed through the authenticated API endpoint using the filetoken.
* **File Token Based:** Uses a unique filetoken identifier to access files stored in PDF.co's built-in file storage.
Files must be uploaded to PDF.co's built-in file storage at [https://app.pdf.co/files](https://app.pdf.co/files) to obtain a filetoken. The filetoken is used to securely reference and retrieve files through the API.
## Request Headers
| Header | Type | Required | Description |
| ----------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------ |
| `x-api-key` | string | *Yes* | Your API key for authentication. Get your API key by registering at [https://app.pdf.co](https://app.pdf.co) |
## Response
The endpoint returns the file content directly with the appropriate Content-Type header based on the file type. The response is a binary file stream.
### Response Headers
| Header | Type | Description |
| --------------------- | ------- | ---------------------------------------------------------------- |
| `Content-Type` | string | The MIME type of the file (e.g., `application/pdf`, `image/png`) |
| `Content-Disposition` | string | The filename and disposition information |
| `Content-Length` | integer | The size of the file in bytes |
### Error Responses
If an error occurs, the endpoint returns a JSON response with the following structure:
| Parameter | Type | Description |
| ----------- | ------- | ----------------------------------------------------------------------------------------------------------------------- |
| `error` | boolean | Indicates whether an error occurred (`true` means error) |
| `status` | integer | Status code of the request (200, 404, 403, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `message` | string | Error message describing what went wrong |
| `errorCode` | integer | Error code of the request (400, 401, 403, 404, 500, etc.) |
## `Example` Response (Success)
On successful request, the endpoint returns the file binary content directly. The following example shows the response headers for a PDF file:
**Response Headers:**
```
Content-Type: application/pdf
Content-Disposition: attachment; filename="document.pdf"
Content-Length: 245678
```
## `Example` Error Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"status": "error",
"errorCode": 404,
"error": true,
"message": "record not found. IMPORTANT: If you need to set JSON data then convert it into string first (e.g. using JSON.stringify(obj) ). Check https://developer.pdf.co for more details."
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/file/download/YOUR_FILETOKEN_HERE' \
--header 'x-api-key: *******************' \
--output downloaded-file.pdf
```
```javascript theme={null}
var https = require("https");
var fs = require("fs");
const API_KEY = "*************************************";
const FILETOKEN = "YOUR_FILETOKEN_HERE";
function downloadFile(apiKey, filetoken) {
return new Promise((resolve, reject) => {
// Prepare request to `file/download` API endpoint
let reqOptions = {
host: "api.pdf.co",
path: `/v1/file/download/${filetoken}`,
headers: { "x-api-key": apiKey }
};
// Send request
https.get(reqOptions, (response) => {
if (response.statusCode === 200) {
// Create write stream for downloaded file
const fileStream = fs.createWriteStream("downloaded-file.pdf");
response.pipe(fileStream);
fileStream.on("finish", () => {
fileStream.close();
console.log("File downloaded successfully");
resolve();
});
} else {
// Handle error response
let data = "";
response.on("data", (chunk) => {
data += chunk;
});
response.on("end", () => {
try {
const errorData = JSON.parse(data);
console.log("Error: " + errorData.message);
reject(errorData);
} catch (e) {
console.log("Error: " + response.statusCode);
reject(new Error(`HTTP ${response.statusCode}`));
}
});
}
})
.on("error", (e) => {
// Request error
console.log("error: " + e);
reject(e);
});
});
}
downloadFile(API_KEY, FILETOKEN);
```
```python theme={null}
import requests
import os
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# File token from PDF.co file storage
FILETOKEN = "YOUR_FILETOKEN_HERE"
# Destination file path
DESTINATION_FILE = "downloaded-file.pdf"
# Prepare URL for file download
url = f"{BASE_URL}/file/download/{FILETOKEN}"
# Execute request
response = requests.get(url, headers={"x-api-key": API_KEY}, stream=True)
if response.status_code == 200:
# Save file to disk
with open(DESTINATION_FILE, "wb") as file:
for chunk in response.iter_content(chunk_size=8192):
file.write(chunk)
print(f"File downloaded successfully as '{DESTINATION_FILE}'")
else:
# Handle error response
try:
error_data = response.json()
print(f"Error: {error_data.get('message', 'Unknown error')}")
except:
print(f"Error: HTTP {response.status_code}")
```
```php theme={null}
```
```csharp theme={null}
using System;
using System.Net;
using System.IO;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
const string FILETOKEN = "YOUR_FILETOKEN_HERE";
const string DESTINATION_FILE = "downloaded-file.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare URL for file download
string url = $"https://api.pdf.co/v1/file/download/{FILETOKEN}";
try
{
// Download file
webClient.DownloadFile(url, DESTINATION_FILE);
Console.WriteLine($"File downloaded successfully as '{DESTINATION_FILE}'");
}
catch (WebException e)
{
Console.WriteLine($"Error: {e.Message}");
}
finally
{
webClient.Dispose();
}
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
# Delete Temporary File
Source: https://developer.pdf.co/api/file-upload/delete
Deletes temporary file (that was uploaded by you or generated by API).
**Try it live:** [Delete Temporary File → API Tester](/api-tester/file-upload/delete) — send a real request from your browser.
## `POST /v1/file/delete`
All temporary files are auto removed after 1 hour. You may use [File Upload](/api/file-upload/overview#temporary-files-upload) methods to explicitly force remove temp files once you don't need them.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| --------- | ------ | -------- | ------- | -------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL of the previously uploaded temporary file or output file that was generated by the API method. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `message` | string | Message of the request |
| `credits` | integer | Number of credits consumed by the request |
| `duration` | integer | Time taken for the operation in milliseconds |
| `errorCode` | integer | Error code of the request (400, 401, 402, 403, 404, 500, etc.) |
| `remainingCredits` | integer | Number of credits remaining in the account |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"error": false,
"status": 200,
"remainingCredits": 9999986
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/file/delete'
--header 'x-api-key: *******************'
--data-raw '{
"url": "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
const API_KEY = "*************************************";
function deleteFile(apiKey, fileName) {
return new Promise(resolve => {
// Prepare request to `file/delete` API endpoint
let queryPath = `/v1/file/delete?url=${fileName}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": apiKey }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.status == 200) {
console.log("remainingCredits: " + data.remainingCredits);
resolve([data.remainingCredits]);
}
else {
// Service reported error
console.log("Error");
}
});
})
.on("error", (e) => {
// Request error
console.log("error: " + e);
});
});
}
let result = deleteFile(API_KEY, "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf");
```
```python theme={null}
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
fileName = "https://pdf-temp-files.s3.amazonaws.com/b5c1e67d98ab438292ff1fea0c7cdc9d/sample.pdf"
url = "{}/file/delete?url={}".format(BASE_URL, fileName)
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["status"] == 200:
remainingCredits = json["remainingCredits"]
```
```php theme={null}
```
# Generate Pre-signed URL
Source: https://developer.pdf.co/api/file-upload/generate-presigned-url
Generate a presigned URL you can PUT a local file to, returning an accessible link for use with other PDF.co API endpoints.
**Try it live:** [Generate Pre-signed URL → API Tester](/api-tester/file-upload/generate-presigned-url) — send a real request from your browser.
## `GET /v1/file/upload/get-presigned-url`
With this method you can upload files up to 2GB in size. Please note that to process these files you should use async=true mode with data extraction and tools endpoints along with [Job Check](/api/job-check) to check status of background jobs you create.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `presignedUrl` | string | The presigned URL to upload the file |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"presignedUrl": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/A1VGV42YE0NWXMKEB4BUIWNYGKXEWTND/test.pdf?X-Amz-Expires=900&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAIZJDPLX6D7EHVCKA/20220913/us-west-2/s3/aws4_request&X-Amz-Date=20220913T074159Z&X-Amz-SignedHeaders=content-type;host&X-Amz-Signature=53f326afde5bcfb3b2714ee8cb5322795bf10a03feb7dab3764e6ca63c017f43",
"url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/A1VGV42YE0NWXMKEB4BUIWNYGKXEWTND/test.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEBgaDLZTUxFLOwF9iiGk%2FyKCATiLp%2FRn9nPmt%2Fey9PcilcRMXtLl0TS6IFNOpk%2BKtSF%2B%2BEVcbNFThw4c1KVx21RQxT5zf7csSEESGov1Xd4uDhF0xGoVkXff9saXGVUtgKrYgPKhUfv5KEO7gz3E0t%2FqCPZJn2KGs1yMbUkohzeIrEd0NH8EVvqfxrfCcW0ZANiG2iMoh8eAmQYyKLjRMfg02ZJPTgoFPQmfMyYt0FacTg4RhkP3PeD9mrWLefDXCwcYkkI%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHFHQYL4OV/20220913/us-west-2/s3/aws4_request&X-Amz-Date=20220913T074159Z&X-Amz-SignedHeaders=host&X-Amz-Signature=9b1a90f36635459bb40f09b0fc6fe3eba185ba3cfdb0a8ef1096ac9efa9b6299",
"error": false,
"status": 200,
"name": "test.pdf",
"credits": 7,
"duration": 0,
"remainingCredits": 98191146
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET https://api.pdf.co/v1/file/upload/get-presigned-url?name=test.pdf&encrypt=true
--header 'x-api-key: YOUR_API_KEY'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
const API_KEY = "**********************************************";
function getPresignedUrl(apiKey) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?name=test.pdf`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": apiKey }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
console.log("presignedUrl: " + data.presignedUrl);
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
let result = getPresignedUrl(API_KEY);
```
```python theme={null}
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
fileName = "test.pdf"
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
```
```php theme={null}
```
# Get MD5 Hash of File by URL
Source: https://developer.pdf.co/api/file-upload/hash
Calculate the MD5 hash of a file from its URL, useful for detecting whether a source document has been modified between requests.
**Try it live:** [Get MD5 Hash of File by URL → API Tester](/api-tester/file-upload/hash) — send a real request from your browser.
## `POST /v1/file/hash`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| --------- | ------ | -------- | ------- | -------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | -------------------------------------------- |
| `hash` | string | Hash of the final PDF file stored in S3. |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"hash": "d942e5becdcb0386598cce15e9e56deb1ca9d893b8578a88eca4a62f02c4000b",
"remainingCredits": 98143
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/file/hash'
--header 'x-api-key: *******************'
--data-raw '{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf"
}'
```
# File Upload Overview
Source: https://developer.pdf.co/api/file-upload/overview
You can upload files as temporary files into PDF.co. Temporary files are stored for 1 hour by default and then auto removed.
To store files permanently (pdf templates, images you want to reuse) please use [PDF.co Built-In Files Storage](https://app.pdf.co/files) instead.
You can also use 3rd party cloud services:
* **Dropbox**: you can use `public link` to a file from Dropbox.
* **Google Drive**: you can use link to a file that was shared as `anyone with a link`.
* **Google Docs/Sheets/Slides**: you can use a link to a document in Google Docs that was shared as `anyone with a link`.
* Any publicly accessible URL from any cloud service or web source that provides a direct link to the uploaded file.
**IMPORTANT NOTE FOR GOOGLE DRIVE/DOCS** users: free Google Drive/Docs limits the number of requests to their files. If you use a link to file or document from Google Drive or Google Drive then make sure you have no more than 5-10 requests per minute. Otherwise Google Drive returns no file or error page.
## Temporary Files Upload
You can upload temporary files up to 2GB in size. Please note that to process these files you should use `async=true` mode with data extraction and tools endpoints along with [/job/check](/api/job-check) to check status of background jobs you create.
## Steps to Upload File
1. First, call [/file/upload/get-presigned-url](/api/file-upload/generate-presigned-url). It will generate link for uploading (`presignedUrl`) and final link (`url`).
2. Now send your file to the `presignedUrl` link using the `PUT` method within the next 30 minutes.
3. Once finished, use `url` to access the file you have just uploaded.
Note: all uploaded files are considered to be temporary files and are automatically permanently removed after 1 hour.
# Upload Small File
Source: https://developer.pdf.co/api/file-upload/upload
Uploads a small (up to 100KB) local file as a temporary file in PDF.co storage. Note: temporary files are automatically permanently removed after 1 hour.
**Try it live:** [Upload Small File → API Tester](/api-tester/file-upload/upload) — send a real request from your browser.
## `POST /v1/file/upload`
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/1a4a92ac805c41c28ef75a24e0f35ba5/sample.pdf",
"error": false,
"status": 200,
"name": "sample.pdf",
"remainingCredits": 98145
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/file/upload'
--header 'x-api-key: *******************'
--form 'file=@"/path/to/file"'
```
```javascript theme={null}
var request = require('request');
var fs = require('fs');
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/file/upload',
'headers': {
'x-api-key': '{{x-api-key}}'
},
formData: {
'file': {
'value': fs.createReadStream('/path/to/file'),
'options': {
'filename': 'filename'
'contentType': null
}
}
}
};
request(options, function (error, response) {
if (error) throw new Error(error);
let data = JSON.parse(response.body);
console.log(data);
});
```
```python theme={null}
import requests
url = "https://api.pdf.co/v1/file/upload"
payload = {}
files = [
('file', open('/path/to/file','rb'))
]
headers = {
'x-api-key': '{{x-api-key}}'
}
response = requests.request("POST", url, headers=headers, json = payload, files = files)
print(response.text.encode('utf8'))
```
```php theme={null}
"https://api.pdf.co/v1/file/upload",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "POST",
CURLOPT_POSTFIELDS => array('file'=> new CURLFILE('/path/to/file')),
CURLOPT_HTTPHEADER => array(
"x-api-key: {{x-api-key}}"
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
?>
```
```csharp theme={null}
using System;
using RestSharp;
namespace HelloWorldApplication {
class HelloWorld {
static void Main(string[] args) {
var client = new RestClient("https://api.pdf.co/v1/file/upload");
client.Timeout = -1;
var request = new RestRequest(Method.POST);
request.AddHeader("x-api-key", "{{x-api-key}}");
request.AddFile("file", "/path/to/file");
IRestResponse response = client.Execute(request);
Console.WriteLine(response.Content);
}
}
}
```
# Upload File Using Base64
Source: https://developer.pdf.co/api/file-upload/upload-base64
Upload a file as base64-encoded source data and receive a temporary URL; the file is automatically removed after one hour.
**Try it live:** [Upload File Using Base64 → API Tester](/api-tester/file-upload/upload-base64) — send a real request from your browser.
## `POST /v1/file/upload/base64`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------ | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `file` | string | *Yes* | - | Base64-encoded file bytes. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/7588d614c9ad41eb98ec317a02abda63/uploadfile.txt",
"error": false,
"status": 200,
"remainingCredits": 77769
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/file/upload/base64'
--header 'x-api-key: *******************'
--form 'file="data:image/gif;base64,R0lGODlhEAAQAPUtACIiIScnJigoJywsLDIyMjMzMzU1NTc3Nzg4ODk5OTs7Ozw8PEJCQlBQUFRUVFVVVVhYWG1tbXt7fInDRYvESYzFSo/HT5LJVJPJVJTKV5XKWJbKWZbLWpfLW5jLXJrMYaLRbaTScKXScKXScafTdIGBgYODg6alprLYhbvekr3elr3el9Dotf///wAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACH5BAAAAAAAIf8LSW1hZ2VNYWdpY2sNZ2FtbWE9MC40NTQ1NQAsAAAAABAAEAAABpJAFGgkKhpFIRHpw2qBLJiLdCrNTFKt0wjD2Xi/G09l1ZIwRJeNZs3uUFQtEwCCVrM1bnhJYHDU73ktJQELBH5pbW+CAQoIhn94ioMKB46HaoGTB5WPaZmMm5wOIRcekqChliIZFXqoqYYkE2SaoZuWH1gmAgsIvr8ICQUPTRIABgTJyskFAw1ZDBAO09TUDw0RQQA7"'
```
# Upload File via Pre-signed URL
Source: https://developer.pdf.co/api/file-upload/upload-presigned-url-put
Upload files up to 100MB directly via HTTP PUT to a presigned URL, for use with async-mode data extraction and processing endpoints.
`PUT {presigned url}`
**Important** The presigned URL must be retreived from the /file/upload/get-presigned-url for the PUT operation to succeed.
Content-Type header
When sending PUT request don't forget to add Content-Type header with proper value based on input file type.
For example:
| File Extension | Content-Type Value |
| ---------------------- | -------------------------- |
| `.txt .csv .xml .json` | text/plain |
| `.pdf` | application/pdf |
| `.msg .eml` | application/vnd.ms-outlook |
| `.doc` | application/msword |
If you're not sure then use application/octet-stream header. It works for most file types.
All uploaded files are treated as temporary files and are automatically permanently removed after 1 hour. If you have a file that you want to reuse over and over, please upload it to PDF.co Built-In Files Storage and get its filetoken:// link that you may reuse inside PDF.co API.
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"presignedUrl": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c72bf56341142ba83c8f98b47f14d62/test.pdf?X-Amz-Expires=900&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAIZJDPLX6D7EHVCKA/20200302/us-west-2/s3/aws4_request&X-Amz-Date=20200302T143951Z&X-Amz-SignedHeaders=host&X-Amz-Signature=8650913644b6425ba8d52b78634698e5fc8970157d971a96f0279a64f4ba87fc",
"url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c72bf56341142ba83c8f98b47f14d62/test.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEGgaDA9KaTOXRjkCdCqSTCKBAW9tReCLk1fVTZBH9exl9VIbP8Gfp1pE9hg6et94IBpNamOaBJ6%2B9Vsa5zxfiddlgA%2BxQ4tpd9gprFAxMzjN7UtjU%2B2gf%2FKbUKc2lfV18D2wXKd1FEhC6kkGJVL5UaoFONG%2Fw2jXfLxe3nCfquMEDo12XzcqIQtNFWXjKPWBkQEvmii4tfTyBTIot4Na%2BAUqkLshH0R7HVKlEBV8btqa0ctBjwzwpWkoU%2BF%2BCtnm8Lm4Eg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHEGHTOA4W/20200302/us-west-2/s3/aws4_request&X-Amz-Date=20200302T143951Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=243419ac4a9a315eebc2db72df0817de6a261a684482bbc897f0e7bb5d202bb9",
"error": false,
"status": 200,
"name": "test.pdf",
"remainingCredits": 98145
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request PUT ''
--header 'x-api-key: YOUR_API_KEY'
--header 'Content-Type: application/octet-stream'
--data-binary '@./sample.pdf'
```
```javascript theme={null}
function uploadFile(apiKey, localFile, uploadFileUrl) {
return new Promise(resolve => {
fs.readFile(localFile, (err, data) => {
request({
method: "PUT",
url: uploadFileUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream",
"x-api-key": apiKey
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + err);
}
});
});
});
}
```
```python theme={null}
uploadFileUrl = "file URL retrieved from /file/upload/get-presigned-url"
with open(fileName, 'rb') as file:
requests.put(uploadFileUrl, data=file, headers={"x-api-key": API_KEY, "content-type": "application/octet-stream"})
```
```php theme={null}
```
# Upload File from URL [GET]
Source: https://developer.pdf.co/api/file-upload/upload-url-get
Download a file from a source URL using a GET request and store it as a temporary PDF.co file, auto-deleted after one hour.
**Try it live:** [Upload File from URL → API Tester](/api-tester/file-upload/upload-url-get) — send a real request from your browser.
## `GET /v1/file/upload/url`
This method do is same as /v1/file/upload/url but using get method.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/97415d1c45a04b29ac42c8dc01883316/sample.pdf",
"error": false,
"status": 200,
"name": "sample.pdf",
"remainingCredits": 77765
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/file/upload/url?url=pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf'
--header 'x-api-key: ******************'
```
# Upload File from URL [POST]
Source: https://developer.pdf.co/api/file-upload/upload-url-post
Download a file from a source URL using a POST request and store it as a temporary PDF.co file, auto-deleted after one hour.
**Try it live:** [Upload File from URL → API Tester](/api-tester/file-upload/upload-url-post) — send a real request from your browser.
## `POST /v1/file/upload/url`
This method do is same as /v1/file/upload/url but using post method.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/1a4a92ac805c41c28ef75a24e0f35ba5/sample.pdf",
"error": false,
"status": 200,
"name": "sample.pdf",
"remainingCredits": 98145
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/file/upload'
--header 'x-api-key: *******************'
--form 'file=@"/path/to/file"'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
const API_KEY = "*************************************";
function upload(apiKey, fileName) {
return new Promise(resolve => {
// Prepare request to `file/upload/url` API endpoint
let queryPath = `/v1/file/upload/url?url=${fileName}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": apiKey }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.status == 200) {
console.log("temp url: " + data.url);
console.log("remainingCredits: " + data.remainingCredits);
resolve([data.remainingCredits]);
}
else {
// Service reported error
console.log("Error");
}
});
})
.on("error", (e) => {
// Request error
console.log("error: " + e);
});
});
}
let result = upload(API_KEY, "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf");
```
```python theme={null}
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
fileName = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"
url = "{}/file/upload/url?url={}".format(BASE_URL, fileName)
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["status"] == 200:
temp_url = json["url"]
remainingCredits = json["remainingCredits"]
```
```php theme={null}
```
# PDF Forms Info Reader
Source: https://developer.pdf.co/api/forms/info-reader
Get information about fillable form fields inside a **PDF** file.
**Try it live:** [PDF Forms Info Reader → API Tester](/api-tester/forms/info-reader) — send a real request from your browser.
## `POST /v1/pdf/info/fields`
For one-time check of PDF file information and find form fields please use PDF [Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper).
Extracts information about fillable PDF fields (fillable edit boxes, fillable check-boxes, radio buttons, combo-boxes) from input PDF file along with general information about the input PDF document. The purpose of this endpoint is to get information about fillable PDFs for use with PDF.co [PDF Add](/api/pdf-add) method.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------- | ------ | ------------- |
| `info` | object | Info details. |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"info": {
"PageCount": 3,
"Author": "SE:W:CAR:MP",
"Title": "2019 Form 1040",
"Producer": "macOS Version 10.15.1 (Build 19B88) Quartz PDFContext",
"Subject": "U.S. Individual Income Tax Return",
"CreationDate": "8/7/2020 11:17:29 AM",
"Bookmarks": "",
"Keywords": "Fillable",
"Creator": "Adobe LiveCycle Designer ES 9.0",
"Encrypted": false,
"PasswordProtected": false,
"PageRectangle": {
"Location": {
"IsEmpty": true,
"X": 0,
"Y": 0
},
"Size": "612, 792",
"X": 0,
"Y": 0,
"Width": 612,
"Height": 792,
"Left": 0,
"Top": 0,
"Right": 612,
"Bottom": 792,
"IsEmpty": false
},
"ModificationDate": "8/7/2020 11:17:29 AM",
"EncryptionAlgorithm": "None",
"PermissionPrinting": true,
"PermissionModifyDocument": true,
"PermissionContentExtraction": true,
"PermissionModifyAnnotations": true,
"PermissionFillForms": true,
"PermissionAccessibility": true,
"PermissionAssemble": true,
"PermissionHighQualityPrint": true,
"FieldsInfo": {
"Fields": [
{
"PageIndex": 1,
"Type": "CheckBox",
"FieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[3]",
"Value": "False",
"Left": 340.39898681640625,
"Top": 67.99798583984375,
"Width": 8,
"Height": 8
},
{
"PageIndex": 1,
"Type": "CheckBox",
"FieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[4]",
"Value": "False",
"Left": 441.1990051269531,
"Top": 67.99798583984375,
"Width": 8,
"Height": 8
},
{
"PageIndex": 1,
"Type": "EditBox",
"FieldName": "topmostSubform[0].Page1[0].f1_03[0]",
"Value": "",
"Left": 238.60000610351562,
"Top": 111.9990234375,
"Width": 228.39999389648438,
"Height": 14.0009765625
},
{
"PageIndex": 2,
"Type": "EditBox",
"FieldName": "topmostSubform[0].Page2[0].PaidPreparer[0].Preparer[0].f2_37[0]",
"Value": "",
"Left": 509.7449951171875,
"Top": 474.0010070800781,
"Width": 66.2550048828125,
"Height": 11.998992919921875
}
]
}
},
"error": false,
"status": 200,
"remainingCredits": 59987
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/info/fields' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"async": false
}'
```
```javascript theme={null}
var request = require('request');
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/pdf/info/fields',
'headers': {
'x-api-key': '{{x-api-key}}'
},
formData: {
'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf'
}
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file url. You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdf-temp-files.s3.amazonaws.com/R2FBM39LFX1BFC860O06XU0TL613JTZ9/f1040-form-filled.pdf "
Async = "False"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
parameters = {}
parameters["async"] = Async
parameters["name"] = os.path.basename(DestinationFile)
parameters["url"] = SourceFileURL
# Prepare URL for 'Info Fields' API request
url = "{}/pdf/info/fields".format(BASE_URL)
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
for field in json["info"]["FieldsInfo"]["Fields"]:print(field["FieldName"] + "=>" + field["Value"])
```
```csharp theme={null}
using System;
using RestSharp;
namespace HelloWorldApplication {
class HelloWorld {
static void Main(string[] args) {
var client = new RestClient("https://api.pdf.co/v1/pdf/info/fields");
client.Timeout = -1;
var request = new RestRequest(Method.POST);
request.AddHeader("x-api-key", "{{x-api-key}}");
request.AlwaysMultipartFormData = true;
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
request.AddParameter("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf");
IRestResponse response = client.Execute(request);
Console.WriteLine(response.Content);
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("text/plain");
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM)
.addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf")
.build();
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/info/fields")
.method("POST", body)
.addHeader("x-api-key", "{{x-api-key}}")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
todo
"https://api.pdf.co/v1/pdf/info/fields",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "POST",
CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf'),
CURLOPT_HTTPHEADER => array(
"x-api-key: {{x-api-key}}"
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
var request = require('request');
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/pdf/info/fields',
'headers': {
'x-api-key': '{{x-api-key}}'
},
formData: {
'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf'
}
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
# Getting Started
Source: https://developer.pdf.co/api/index
Introducing the general concepts for using the PDF.co API, authentication methods, response codes and sample code.
## API Reference
The PDF.co Web API is REST-based, making it intuitive and easy to use. To prioritize your data’s security and privacy, all requests are securely transmitted using HTTPS. Kindly note, unsecured HTTP connections are not supported.
All requests contain the following **base URL**:
`https://api.pdf.co/v1`
## Authenticating Your API Request
To authenticate you need to add a header named `x-api-key` using your API Key as the value.
```javascript theme={null}
"x-api-key": "sample@sample.com_123a4b567c890d123e456f789g01"
```
The key provided above is just a sample and won’t work for actual API calls. Don’t forget to replace it with your real API Key, which you can find in your [PDF.co Dashboard](https://app.pdf.co/), when making requests.
## Response codes
After making a request you will receive a response from the **PDF.co** API. A code `200` means the request was successfull, a `400` means there was an error. However there could be other codes - see [the complete list of available response codes](/api/response-codes).
## Sample Code
Here is some sample code which would convert a **PDF** to a **CSV** file.
```javascript theme={null}
var data = {
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf",
"lang": "eng",
"inline": true,
"pages": "0-",
"async": false,
"name": "result.csv"
}
fetch('https://api.pdf.co/v1/pdf/convert/to/csv', {
method: 'POST',
headers: {
'Accept': 'application/json',
'Content-Type': 'application/json',
'x-api-key': 'sample@sample.com_123a4b567c890d123e456f789g01'
},
body: JSON.stringify(data)
})
.then(response => response.json())
.then(response => console.log(JSON.stringify(response)))
```
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/csv' \
--header 'Content-Type: application/json' \
--header 'x-api-key: sample@sample.com_123a4b567c890d123e456f789g01' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf",
"lang": "eng",
"inline": true,
"pages": "0-",
"async": false,
"name": "result.csv"
}'
```
# Background & Job Check
Source: https://developer.pdf.co/api/job-check
Checks the [status](#available-status-values) of a background job that was previously created with PDF.co API. Use this API to check the status of your asynchronous API calls.
**Try it live:** [Background & Job Check → API Tester](/api-tester/job-check) — send a real request from your browser.
## `POST /v1/job/check`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| --------- | ------- | -------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `jobId` | string | *Yes* | - | ID of background that was started asynchronously. To start a new async background job, you should set async to true for API methods. |
| `force` | boolean | *No* | `false` | Set to true to forcibly check the status of the background job. Intended to be used with really long and heavy background jobs only. |
### Available Status Values
* `working` - background job is currently in work or does not exist.
* `success` - background job was successfully finished.
* `failed` - background job failed for some reason (see `message` for more details).
* `aborted` - background job was aborted.
* `unknown` - unknown background job id. Available only when force is set to `true` for input request.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `message` | string | Message of the request |
| `pageCount` | integer | Number of pages in the PDF document. |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `jobId` | string | Identifier for the job request |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `jobDuration` | integer | Time taken to execute the job in milliseconds |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"jobid": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. 1
```json theme={null}
{
"status": "working",
"remainingCredits": 60227
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page. 2
```json theme={null}
{
"status": "success",
"message": "Success",
"url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300/L8QYIZQ6KZOITCT0PXUNPM6HKYSP5OIO.json?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzECcaDAbrXwAd1IYG3nZR5yKCAdcavWT%2BuwTotGsad9asqRzowPa1M4BoIWU0M9FqXNJP8xBIQX1Cn7XTq4ZfpklsxcpGE4WcapfHdooi2uR1QWw4kuUlMGGU92uy7pS0RhaGCEL00ES%2BIb%2F5039yyAFklqfAgDlHvi47I7Pp01y6Ua25RzrZGh6ACOd7le%2BXArnbQs4o4ezNqgYyKD%2FCX1I5ZOS0tu0ND0I%2FUWTHp6OR8He9a0dgVXfiMU7pNkwQqwVVFcM%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAZTLLKK5/20231114/us-west-2/s3/aws4_request&X-Amz-Date=20231114T134932Z&X-Amz-SignedHeaders=host&X-Amz-Signature=e5553e080a23fb158c0514f99c9f70be0cb74f764933d712ba628110d4079b4c",
"jobId": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300",
"credits": 2,
"remainingCredits": 1480582,
"duration": 33
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/job/check' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"jobid": "6YSZD3U872ZYYFEDMQCQSGEEO8YSF5WA--151-300"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
const jobId = "{your_job_id}";
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "success") {
console.log(`Job success!`);
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
jobId = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
// Prepare POST request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("jobid", new JsonPrimitive(jobId));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request to `Job Check` API
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/job/check")
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
}
}
```
```php theme={null}
function CheckJobStatus($jobId, $apiKey)
{
$status = null;
// Create URL
$url = "https://api.pdf.co/v1/job/check";
// Prepare requests params
$parameters = array();
$parameters["jobid"] = $jobId;
// Create Json payload
$data = json_encode($parameters);
// Create request
$curl = curl_init();
curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey, "Content-type: application/json"));
curl_setopt($curl, CURLOPT_URL, $url);
curl_setopt($curl, CURLOPT_POST, true);
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($curl, CURLOPT_POSTFIELDS, $data);
// Execute request
$result = curl_exec($curl);
if (curl_errno($curl) == 0)
{
$status_code = curl_getinfo($curl, CURLINFO_HTTP_CODE);
if ($status_code == 200)
{
$json = json_decode($result, true);
if (!isset($json["error"]) || $json["error"] == false)
{
$status = $json["status"];
}
else
{
// Display service reported error
echo "
Error: " . $json["message"] . "
";
}
}
else
{
// Display request error
echo "
Status code: " . $status_code . "
";
echo "
" . $result . "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Language Support
Source: https://developer.pdf.co/api/language-support
Reference list of OCR language codes supported by PDF.co for text extraction and searchable PDF generation.
## Supported Language Codes
| Code | Description |
| ---------- | ------------------------------ |
| `afr` | Afrikaans |
| `amh` | Amharic |
| `ara` | Arabic |
| `asm` | Assamese |
| `aze` | Azerbaijani |
| `aze_cyrl` | Azerbaijani - Cyrillic |
| `bel` | Belarusian |
| `ben` | Bengali |
| `bod` | Tibetan |
| `bos` | Bosnian |
| `bul` | Bulgarian |
| `cat` | Catalan; Valencian |
| `ceb` | Cebuano |
| `ces` | Czech |
| `chi_sim` | Chinese - Simplified |
| `chi_tra` | Chinese - Traditional |
| `chr` | Cherokee |
| `cym` | Welsh |
| `dan` | Danish |
| `deu` | German |
| `dzo` | Dzongkha |
| `ell` | Greek, Modern (1453-) |
| `eng` | English |
| `enm` | English, Middle (1100–1500) |
| `epo` | Esperanto |
| `est` | Estonian |
| `eus` | Basque |
| `fas` | Persian |
| `fin` | Finnish |
| `fra` | French |
| `frk` | Frankish |
| `frm` | French, Middle (ca. 1400–1600) |
| `gle` | Irish |
| `glg` | Galician |
| `grc` | Greek, Ancient (-1453) |
| `guj` | Gujarati |
| `hat` | Haitian; Haitian Creole |
| `heb` | Hebrew |
| `hin` | Hindi |
| `hrv` | Croatian |
| `hun` | Hungarian |
| `iku` | Inuktitut |
| `ind` | Indonesian |
| `isl` | Icelandic |
| `ita` | Italian |
| `ita_old` | Italian - Old |
| `jav` | Javanese |
| `jpn` | Japanese |
| `kan` | Kannada |
| `kat` | Georgian |
| `kat_old` | Georgian - Old |
| `kaz` | Kazakh |
| `khm` | Central Khmer |
| `kir` | Kirghiz; Kyrgyz |
| `kor` | Korean |
| `kur` | Kurdish |
| `lao` | Lao |
| `lat` | Latin |
| `lav` | Latvian |
| `lit` | Lithuanian |
| `mal` | Malayalam |
| `mar` | Marathi |
| `mkd` | Macedonian |
| `mlt` | Maltese |
| `msa` | Malay |
| `mya` | Burmese |
| `nep` | Nepali |
| `nld` | Dutch; Flemish |
| `nor` | Norwegian |
| `ori` | Oriya |
| `pan` | Panjabi; Punjabi |
| `pol` | Polish |
| `por` | Portuguese |
| `pus` | Pushto; Pashto |
| `ron` | Romanian; Moldavian; Moldovan |
| `rus` | Russian |
| `san` | Sanskrit |
| `sin` | Sinhala; Sinhalese |
| `slk` | Slovak |
| `slv` | Slovenian |
| `spa` | Spanish; Castilian |
| `spa_old` | Spanish; Castilian - Old |
| `sqi` | Albanian |
| `srp` | Serbian |
| `srp_latn` | Serbian - Latin |
| `swa` | Swahili |
| `swe` | Swedish |
| `syr` | Syriac |
| `tam` | Tamil |
| `tel` | Telugu |
| `tgk` | Tajik |
| `tgl` | Tagalog |
| `tha` | Thai |
| `tir` | Tigrinya |
| `tur` | Turkish |
| `uig` | Uighur; Uyghur |
| `ukr` | Ukrainian |
| `urd` | Urdu |
| `uzb` | Uzbek |
| `uzb_cyrl` | Uzbek - Cyrillic |
| `vie` | Vietnamese |
| `yid` | Yiddish |
# Merge PDF
Source: https://developer.pdf.co/api/merge/pdf
Merge multiple PDF files into a single PDF document.
**Try it live:** [Merge PDF → API Tester](/api-tester/merge/pdf) — send a real request from your browser.
## `POST /v1/pdf/merge`
The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------------- | ------- | -------- | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `RenameMatchingFieldsDuringMerge` | boolean | *No* | `true` | This feature enables the renaming of field names during the merging of PDF files which contain forms. If set to false, it will retain the original field names. This is helpful for merged PDF forms with identical field names when the customer wants to auto-fill the identical field names in other pages. |
| `GenerateBookmarks` | boolean | *No* | `false` | This adds bookmarks to the merged document with names assigned to every merged document in the same order: |
| `zipIncludeFilter` | string | *No* | - | You can control which files to include and exclude from input zip files with a profiles. |
| `zipExcludeFilter` | string | *No* | - | zipIncludeFilter and zipExcludeFilter support \* and ? wildcards. |
| `MergedDocumentTitle` | string | *No* | Title of the first document | Specifies a custom title for the merged document. Overrides the title of the first document during the merge process. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Rename Matching Fields
This feature enables the renaming of field names during the merging of **PDF** files which contain forms. If set to `false`, it will retain the original field names. This is helpful for merged **PDF** forms with identical field names when the customer wants to auto-fill the identical field names in other pages.
```
{
"profiles": "{ 'RenameMatchingFieldsDuringMerge': false }"
}
```
### Generate Bookmarks
This adds bookmarks to the merged document with names assigned to every merged document in the same order:
```
{
"profiles": "{'GenerateBookmarks': true, 'BookmarkTitles': [ 'BookmarkName1', 'BookmarkName2', 'BookmarkName3' ] }"
}
```
### Include / Exclude from ZIPS
You can control which files to include and exclude from input zip files with a `profiles`.
```json theme={null}
// include PDF, XLS and XLSX files
{
"profiles": "{ 'zipIncludeFilter': '*.pdf,*.xls*' }"
}
```
```json theme={null}
// exclude DOC, DOCX, XLS and XLSX files
{
"profiles": "{ 'zipExcludeFilter': '*.doc*,*.xls*' }"
}
```
`zipIncludeFilter` and `zipExcludeFilter` support `*` and `?` wildcards.
### Change Document Title
You can chnage the document title during a merge with the following:
```json theme={null}
{
"profiles": "{ 'MergedDocumentTitle': 'New Title' }"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/3ec287356c0b4e02b5231354f94086f2/result.pdf",
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 98465
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/merge' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URLs of PDF files to merge
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFiles = [
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf"
];
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Merge PDF` API endpoint
var queryPath = `/v1/pdf/merge`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceFiles.join(",")
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "**********************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF files
SourceFile_1 = ".\\sample1.pdf"
SourceFile_2 = ".\\sample2.pdf"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
UploadedFileUrl_1 = uploadFile(SourceFile_1)
UploadedFileUrl_2 = uploadFile(SourceFile_2)
if (UploadedFileUrl_1 != None and UploadedFileUrl_2!= None):
uploadedFileUrls = "{},{}".format(UploadedFileUrl_1, UploadedFileUrl_2)
mergeFiles(uploadedFileUrls, DestinationFile)
def mergeFiles(uploadedFileUrls, destinationFile):
"""Perform Merge using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrls
# Prepare URL for 'Merge PDF' API request
url = "{}/pdf/merge".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URLs of PDF files to merge
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
static string[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf" };
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare URL for `Merge PDF` API call
string url = "https://api.pdf.co/v1/pdf/merge";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", string.Join(",", SourceFiles));
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URLs of PDF files to merge
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample2.pdf" };
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Merge PDF` API call
String query = "https://api.pdf.co/v1/pdf/merge";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
String.join(",", SourceFiles));
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "PDF Merging" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Merge Various Document Type
Source: https://developer.pdf.co/api/merge/various-files
Merge PDF from two or more PDF, DOC, XLS, images, even ZIP with documents and images into a new PDF.
**Try it live:** [Merge Various Document Type → API Tester](/api-tester/merge/various-files) — send a real request from your browser.
## `POST /v1/pdf/merge2`
We do not support images in the HEIC format (Apple’s image format) or the WEBP format.The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed.
This endpoint is similar to [/pdf/merge](/api/merge/pdf) but it also supports `zip`, `doc`, `docx`, `xls`, `xlsx`, `rtf`, `txt`, `png`, `jpg` files as source.
We recommended to use this endpoint in `async: true` mode because it may need to convert source documents to PDF.
This endpoint also consumes more credits because of the internal conversions.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls, https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/images-and-documents.zip",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/3ec287356c0b4e02b5231354f94086f2/result.pdf",
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 98465
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/merge2' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf,https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls, https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/images-and-documents.zip",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URLs of files to merge. Supports documents, spreadsheets, images as sources.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFiles = [
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg"
];
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Merge Documents` API endpoint
var queryPath = `/v1/pdf/merge2`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceFiles.join(","), async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
function checkIfJobIsCompleted(jobId, resultFileUrl) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl); }, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source Document files. Supports documents, spreadsheets, images as sources.
SourceFile_1 = ".\\sample1.pdf"
SourceFile_2 = ".\\sample.docx"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
# (!) Make asynchronous job
Async = True
def main(args = None):
UploadedFileUrl_1 = uploadFile(SourceFile_1)
UploadedFileUrl_2 = uploadFile(SourceFile_2)
if (UploadedFileUrl_1 != None and UploadedFileUrl_2!= None):
uploadedFileUrls = "{},{}".format(UploadedFileUrl_1, UploadedFileUrl_2)
mergeFiles(uploadedFileUrls, DestinationFile)
def mergeFiles(uploadedFileUrls, destinationFile):
"""Perform Merge using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {}
parameters["async"] = Async
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrls
# Prepare URL for 'Merge Document' API request
url = "{}/pdf/merge2".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# URL of the result file
resultFileUrl = json["url"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.IO;
using System.Net;
using Newtonsoft.Json.Linq;
using System.Threading;
using System.Collections.Generic;
using Newtonsoft.Json;
// Cloud API asynchronous "Merge Document" job example. Supports documents, spreadsheets, images as sources.
// Allows to avoid timeout errors when processing huge or scanned PDF documents.
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URLs of PDF files to merge. Supports documents, spreadsheets, images as sources.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
static string[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/other/Input.xls",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg"
};
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
// (!) Make asynchronous job
const bool Async = true;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare URL for `Merge PDF` API call
string url = "https://api.pdf.co/v1/pdf/merge2";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", string.Join(",", SourceFiles));
parameters.Add("async", Async);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Asynchronous job ID
string jobId = json["jobId"].ToString();
// URL of generated PDF file that will available after the job completion
string resultFileUrl = json["url"].ToString();
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do
{
string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status);
if (status == "success")
{
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
break;
}
else if (status == "working")
{
// Pause for a few seconds
Thread.Sleep(3000);
}
else
{
Console.WriteLine(status);
break;
}
}
while (true);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
static string CheckJobStatus(string jobId)
{
using (WebClient webClient = new WebClient())
{
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
string response = webClient.DownloadString(url);
JObject json = JObject.Parse(response);
return Convert.ToString(json["status"]);
}
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URLs of input files to merge. Supports documents, spreadsheets, images as sources.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"
};
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Merge Document` API call
String query = "https://api.pdf.co/v1/pdf/merge2";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
String.join(",", SourceFiles));
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "PDF Merging 2" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF Add
Source: https://developer.pdf.co/api/pdf-add
Add text, images, forms, other PDFs, fill forms, links to external sites and external PDF files. You can update or modify PDF and scanned PDF files.
**Try it live:** [PDF Add → API Tester](/api-tester/pdf-add) — send a real request from your browser.
## `POST /v1/pdf/edit/add`
Create new PDF forms with fillable edit boxes, checkboxes and other fillable fields.
Quickly create configs for [PDF.co](https://pdf.co/) API, [Zapier](https://zapier.com/), [Make with PDF.co PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper).
To add a signature, save the signature as an image and insert it using the images attribute.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `password` | string | *No* | - | Password for the PDF file. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `inline` | boolean | *No* | `false` | Set to `true` to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `annotations` | array\[object] | *No* | - | |
| `text` | string | *Yes* | - | String to add, if you need to insert a line break then use \n or `{{$$newLine}}`. You can also use built-in macros like `{{$$PageNumber}}` and custom data macros. |
| `x` | integer | *Yes* | - | X coordinate (zero point is in the top left corner). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `y` | integer | *Yes* | - | Y coordinate (zero point is in the top left corner). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `type` | string | *No* | `Text` | Set object type. Available types: `text` = text object (default), `textField` = text input field, `TextFieldMultiline` = multiline fillable text field, `checkbox` = checkbox field. |
| `id` | string | *No* | - | Sets id of the form field if type is not text. |
| `leading` | integer | *No* | - | Sets a custom line height for text. The value defines the vertical spacing between lines. Larger values increase the space between lines, while smaller values tighten the spacing. |
| `width` | integer | *No* | - | Width of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. |
| `height` | integer | *No* | - | Height of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. |
| `alignment` | string | *No* | `left` | Sets text alignment within the width of the text box. Valid values: left, center, right. Default is left. |
| `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. |
| `color` | string | *No* | `#000000` | Sets the text color. Default is `#000000` (non-transparent black). Color in **RRGGBB** or **AARRGGBB** format where **AA** is the transparency component. For example, 50% transparent green is `#8000FF00`. |
| `link` | string | *No* | - | Sets link on click for text. |
| `size` | integer | *No* | `12` | Sets font size. |
| `transparent` | boolean | *No* | `true` | Set to `false` to force disable any transparency and draw a white background under the text. |
| `fontName` | string | *No* | `Arial` | Set font name to use. Default is "Arial". See [availabe fonts](#available-fonts). |
| `fontBold` | boolean | *No* | `false` | Set to `true` to enable bold font style. |
| `fontStrikeout` | boolean | *No* | `false` | Set to `true` to enable strikeout font style. |
| `fontUnderline` | boolean | *No* | `false` | Set to `true` to enable underline font style. |
| `RotationAngle` | integer | *No* | `0` | Set rotation angle in degrees. Default is `0` degrees. |
| `images` | array\[object] | *No* | - | |
| `url` | string | *Yes* | - | URL to image or PDF as HTTP link, file token, or datauri:.. URL (with base64 encoded image). |
| `x` | integer | *Yes* | - | X coordinate (zero point is in the top left corner). Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `y` | integer | *Yes* | - | Y coordinate (zero point is in the top left corner). Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. |
| `width` | integer | *No* | - | Width of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. |
| `height` | integer | *No* | - | Height of the text box. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure pdf coordinates. |
| `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. |
| `link` | string | *No* | - | Sets link on click for text. |
| `keepAspectRatio` | boolean | *No* | `true` | Set to `false` if don’t need to keep the aspect ratio for the image or PDF added. In this case, it will use the width and height parameters provided. |
| `fields` | array\[object] | *No* | - | |
| `fieldName` | string | *Yes* | - | Name of the form field. To find form fields please use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). |
| `text` | string | *Yes* | - | Value to set for this field. If you have a checkbox, set X, true, 1, or another text which is different from false to enable the checkbox. For radio buttons and combo boxes, you need to set the item value in text or index of the item to select. To find form fields please use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper). |
| `pages` | string | *No* | - | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0, Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). To process all pages, use "0-". If not specified, the default configuration processes all pages. The input must be in string format. |
| `size` | integer | *No* | - | Override the font size of the text inside the given input field. |
| `fontName` | string | *No* | - | Name of the font to use to fill out the input field. |
| `fontBold` | boolean | *No* | - | Override font bold style of the text input field. |
| `fontItalic` | boolean | *No* | - | Override font italic style of the text input field. |
| `fontStrikeout` | boolean | *No* | - | Override font strikeout style of the text input field. |
| `fontUnderline` | boolean | *No* | - | Override font underline style of the text input field. |
| `annotationsString` | string | *No* | - | This parameter represents one or more text objects to add to a PDF. Each object is made of parameter separated by the `;` symbol. |
| `imagesString` | string | *No* | - | Adds one or more images or other PDF objects on top of the source PDF. Each object is made of parameter separated by the `;` symbol. |
| [`fieldsString`](#fieldsstring) | string | *No* | - | Set values for fillable PDF field objects. Each object is made of parameter separated by the `;` symbol. See [fieldsString](#fieldsstring) for more information. |
| `profiles` | object | *No* | - | - |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `Pages[0].SetCropBox()` | array\[string] | *No* | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | *No* | `false` | To disable ligaturization, for example for Hebrew, use the following: |
| `FlattenDocument()` | boolean | *No* | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### fieldsString
Set values for fillable PDF fields (i.e. fill pdf fields in pdf forms). To fill fields in PDF form, use the following format: `page;fieldName;value`.
Also, the advanced format can be used to override font name, size and style:
`0;fieldName;Field Text;12+bold+italic+underline+strikeout;FontName`
Where:
`0;editbox1;text is here;12+bold;Arial`
To fill the checkbox, use true, for example: `0;checkbox1;true`.
To separate multiple objects, use the `|` separator. To get the list of all fillable fields in a PDF form, use the :ref:post-tag-pdf-info-fields endpoint. If you need to include a pipe character within a value, escape it by prefixing with `\\` (for example, `|` should be written as `\\|` and in a JSON string as `\\|`).
### Crop a PDF File
Crop a **PDF** file using an array to define the crop area. The crop box is defined by a rectangle `[x, y, width, height]` in **PDF** points (1 Point = 1/72 inches).
An A4 page size in points is **595 x 842**
```json theme={null}
{
"profiles": "{ 'Pages[0].SetCropBox()': ['28', '28', '539', '786'] }"
}
```
### Disable Ligaturization
To disable ligaturization, for example for Hebrew, use the following:
```json theme={null}
{
"profiles": "{ 'DisableLigatures': true }"
}
```
### Flatten Document
Flattening a document renders it as *read-only*. Handy if you want to remove editing or copying capability.
```json theme={null}
{
"profiles": "{ 'FlattenDocument()': [] }"
}
```
## Available fonts
### Standard Fonts
* Arial
* Arial Black
* Aptos
* Aptos Display
* Aptos Narrow
* Bahnschrift
* Calibri
* Cambria
* Cambria Math
* Candara
* Comic Sans MS
* Consolas
* Constantia
* Corbel
* Courier New
* Ebrima
* Franklin Gothic Medium
* Gabriola
* Gadugi
* Georgia
* HoloLens MDL2 Assets
* Impact
* Ink Free
* Javanese Text
* Leelawadee UI
* Lucida Console
* Lucida Sans Unicode
* Malgun Gothic
* Marlett
* Microsoft Himalaya
* Microsoft JhengHei
* Microsoft New Tai Lue
* Microsoft PhagsPa
* Microsoft Sans Serif
* Microsoft Tai Le
* Microsoft YaHei
* Montserrat
* Microsoft Yi Baiti
* MingLiU-ExtB
* Mongolian Baiti
* MS Gothic
* MV Boli
* Myanmar Text
* OCR A
* OCR A Extended
* OCR B
* OCR B E
* OCR B F
* OCR B L
* OCR B S
* OCR B X
* Nirmala UI
* Palatino Linotype
* Segoe MDL2 Assets
* Segoe Print
* Segoe Script
* Segoe UI
* Segoe UI Historic
* Segoe UI Emoji
* Segoe UI Symbol
* SimSun
* Sitka Banner
* Sitka Banner Semibold
* Sitka Display
* Sitka Display Semibold
* Sitka Heading
* Sitka Heading Semibold
* Sitka Small
* Sitka Small Semibold
* Sitka Subheading
* Sitka Subheading Semibold
* Sitka Text
* Sitka Text Semibold
* Sylfaen
* Symbol
* Tahoma
* Times New Roman
* Trebuchet MS
* Verdana
* Webdings
* Wingdings
* Yu Gothic
### Japanese Fonts
* MS Gothic
* MS Mincho
* Yu Gothic
### Chinese Fonts
* SimSun
* MingLiU
* Microsoft YaHei
### Korean Fonts
* Malgun Gothic
### Hebrew Fonts
* Miriam
### Arabic Fonts
* Aldhabi
* Andalus
* Arabic Typesetting
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `hash` | string | Hash of the final PDF file stored in S3. |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `pageCount` | integer | Number of pages in the PDF document. |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `credits` | integer | Number of credits consumed by the request |
| `duration` | integer | Time taken for the operation in milliseconds |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
## `Example` Payload (A)
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"async": false,
"inline": true,
"name": "f1040-form-filled",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true",
"imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200",
"fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789"
}
```
## `Example` Response (A)
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/0c336bfcef1a473d98492bda25d8da03/newDocument.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEO7%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDHWK1dY4d4lOgsheliKBATwE%2FZewASPTEnPxTn%2BOdYhP4h3gljAJfqbRvQptDX7wdWLmrBS7Tg4qTU6pAbxIdXChGPjBWpSbtiADJKmqkmyhkUmE8GSM1%2FGtJO6bga2pgzvFLXmzxjTf3%2BFNqwYOvbyApIZdVLoPpEKY6PlCflQtLTd30dhelm6xpB8pitbdhSjdz8KCBjIobVy%2Fjwybwp6OQgB%2FT6QkIo2dU07gtFREdn5jhRyvnS5lkccweBV1%2Bw%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHMV5P3JOS/20210316/us-west-2/s3/aws4_request&X-Amz-Date=20210316T124309Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=95287bf3c007fed4c2c5aeea1ce75c846cc6c68b22aaf35175ebe41a105f54e1",
"pageCount": 1,
"error": false,
"status": 200,
"name": "newDocument",
"remainingCredits": 9913694,
"credits": 3
}
```
## `Example` Payload (B)
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"async": false,
"inline": true,
"name": "newDocument",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf",
"annotations": [
{
"text": "Sample Text 1",
"x": 150,
"y": 100,
"size": 20,
"pages": "0-"
},
{
"text": "sample text that is centered (can also set right or left alignment) ",
"x": "10",
"y": "10",
"width": "500",
"height": "200",
"size": "7",
"pages": "0",
"alignment": "center"
},
{
"text": "Sample Text 2 - Click here to test link\r\n(CLICK ME!)",
"x": 250,
"y": 240,
"size": 24,
"pages": "0-",
"color": "CCBBAA",
"link": "https://pdf.co/",
"fontName": "Comic Sans MS",
"fontItalic": true,
"fontBold": true,
"fontStrikeout": false,
"fontUnderline": true
},
{
"text": "Simple text 3",
"x": 100,
"y": 230,
"size": 12,
"pages": "0-",
"type": "Text"
},
{
"text": "sample text 3 - input text field",
"x": 100,
"y": 170,
"size": 16,
"pages": "0-",
"type": "TextField",
"id": "textfield1"
},
{
"x": 200,
"y": 120,
"size": 16,
"pages": "0-",
"type": "Checkbox",
"id": "checkbox2"
},
{
"x": 200,
"y": 140,
"size": 16,
"pages": "0-",
"type": "CheckboxChecked",
"id": "checkbox3"
}
],
"images": [
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png",
"x": 270,
"y": 150,
"width": 159,
"height": 43,
"pages": "0"
},
{
"url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAgMAAAEtCAYAAACVlWOMAAAgAElEQVR4Xu3dCXxkVZn38f9zK72wiCjdgEInadx1RnFkHDckCQiiIi7grqDMNEmQAXV05p1XBcdlXEFHuxN6FFF5RwUXcGNPAoq7CKKOG3SSRhS6W5ul6Sbpus/7OZWq5FZ1JankJuncvr/6fOYzM6Tuued879N1n3vuWUx8EEAAAQQQQCDXApbr1tN4BBBAAAEEEBDJAEGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEEAAgZwLkAzkPABoPgIIIIAAAiQDxAACCCCAAAI5FyAZyHkA0HwEEEAAAQRIBogBBBBAAAEEci5AMpDzAKD5CCCAAAIIkAwQAwgggAACCORcgGQg5wFA8xFAAAEEECAZIAYQQAABBBDIuQDJQM4DgOYjgAACCCBAMkAMIIAAAgggkHMBkoGcBwDNRwABBBBAgGSAGEAAAQQQQCDnAiQDOQ8Amo8AAggggADJADGAAAIIIIBAzgVIBnIeADQfAQQQQAABkgFiAAEEEJgDgaNPH3piHOlquX66fMnIG6745GPunYNiKQKBBREgGVgQ5t15Erf2M+54ujx+o1zHSDpU0rJyjR6QdLvcbzDTt7YvLwz84PxV23dnbTn3wgscs+a2hxablrxWsb9SpidKWlGuRSzTJne7ydyvtNi+3bd+1e2SeW0tj1wz+Iimgt1o0taouPMl165/1PDCt2T3nrH9jI3/oDi+StLeJj+xr6f1it1bI86OQOMCJAONW2Xum6UnFdMXZPq7ROUflGuTTC6XybSyKjlwfWrfeMl7vrn+kSFR4NOAQPlG+F3JruzvWXVmvZtlA8Us+FdOPvlXS7es2OcdLvs/4QaWqMBmSZWkcK9EchDSgFststOuW9f8k2SFj+38w4Gj0ZLr5Xq8ZGv7e5rfvOAN2s0nTCQDD5X7Sf29rV/dzVXi9Ag0LEAy0DBVtr7Y3jX0NEnfGfshtxvM4/du36vwg9on/3BD2HzAPq+T2UckPbzUStcNvjw+ceDjq7dmq9W7p7YdnRsPcfPvS/FexULh6Bs+terW3VOTxs/6tDU/XfLQwsp1Lr1J0la5f9yXNH124JOH/LE2mWnr3vho8/jjko6XFEl60Mze2Leu+YvJMz63c/CpBelFkv22v7f50qwkRY2rTf1NkoG5kqSc3SFAMrA71Of5nOGHfr/CgV+S/AQ3WzOwbtXnpvthLnUVR00fkZVuDgW5zurvbfmvea7qHlH8RDLgKyPzF163rrV/soa1dW5oNSu8RuZPlOsRkv/Q3f97oHf14EJiHN01fGQsv0KyW6MoetF1aw/dMvX53do7Nx4r83WSDpPpN0vi0aOu7n303QtZ78V8rmPW3NZcjJpulOnQrPYMLJb4XMzXeU+tG8nAHnhljztt48NHl8bXuuwAc3tWX++qPzbazNKPQRS9Qu6/6u9p/XYjNwgzP9+lx0j6o5l9RbGdP5NzNlq3xfq9RDLQPNlNoOOM21s8LqxNPF0nm1M016cO2HL/Oy699EkjC9HO9s7BN8nsM5LO6e9p+Y9Gz1nqSVqxz/NMeqpFTT3TJxGNlrx7vnfCmjv33lYYeavLzpa0r6RfmmztPsWmL8/0VVl1HGQrmV5s8bl7oiHfZyUZ2AOv/3gy4FpZiHc+e74Gc008XWqfGsYHTP6fB2ze9uGFurntzss4XTLQ0TV4vMu+IOkAme6y2D8dR/6VUGfz6K2SXhP+T7mf29/b8r7penHmoq3jyYDZe/vXNb97LsrMYhkdXUMfdOlf69T9dpdOH+hpubbRdj339I2PKUTx9VLo8Zk8yWrvuv1xRY/2vqG39eeNlj2f31uM8Tmf7aXs+gIkA3tkZLh1dA2vL70PnscbTHvXUBh49k5JF0t+vbme6JH9o1wHlVkvHXlw+2k3Xvj4+/ZI5nKjpkoGOroHn+Vu35D0MLk+uuvgTLf27o1d4Z29XHd5FLUPrFv1h/n2auvecLh5FG50fy1Ix13b03L7XJzz+DN/v9+O0SXnWhT9pG/dqi/NV2JTeqJvGn2tXG936VGSRt10Q6Gos6+7oOXXjbSlrfvufS3e/m2ZtUrxp2R2l1wvkPSS8qDaB9111kBv8/pG2lE1ZmCSZKA0/iKO+xXp/vl4zdK25s4VVti5xuSvc+mxpVd+UhgM/EN3+++VW+77WjJBX6zx2cj14ztzK0AyMLeei6a0xFN700x+0GbSgNBlvGnlyqUD6w68v3Jc+JG+Lxp9u5lCohCmMH5+xeb7/2m+eghK3dYr932VvPR01yrZZnP/ehTvPOfa9Y+6Zybtme13q54IE6PISzfGnUvDIM5nTpWUBbP7C6PflNThbq8e6G3+0mzr0uhxiXElL3PTD5ZEdtI1n2q+s9HjJ/teIjHauLxp5AXTzbV/5ls27rX3A/Fz4shXlgYe9jTfNNWNt+6A1+rK3GPub+zrbf16I20JCcHKTZtGqm6QpVc6Ua9kz59ssGS9sstP2JdLWjJZz0B759A/y/SJ0iDdaK8XJv/tTFXf0JtgUdNBfWtXhVkr41M7n/fm4UeOFvX0QuzbYtO/ytRWTgDqF+e6vLIGwmKOz0auHd+ZWwGSgbn1XESlhQFfQ++U2bmlSrm+6VHcObBu9Z/nv5JuHZ3Dr3PTf5d+GOv2Trg9t3Po8CYV7g7jC47u2vjkovm55n5Eub4bFNmXfWfTJQPrHxmmuu3yaesaOsakC0oD2nb93B3JTrqup/m7893eyZ4I27qGTjHpQjf9SEvjF0w1O6O9e/g/5P6umb7DT9O2Y7qGDitKYV78oyX9Ra6uFVuav3rppVacbbmNJgOlAauFpg9LOiUxtTU26SMHbL7/3ZWb8/jTu/yXbvZjk8IrjXC9HzT5Fe76ZuS2MY4U4uZfyjNitpj5i/vWtX5/tu0YG0sweoFLr5NUp7xSj87Zcg8Dbh8q6SGS9p/kfGEcyJ8U2VkW+0nlMr/c39PyqkbqV1nDQdJ+bvExA+tW31w5rtw794Hxclw3mfztUuG3kcXLdro/MzK9yGXPK72mkkYrayAs9vhsxIbvzJ0AycDcWS7CksZvymFaWJg2GLoLP1Eo7vzQ/D81VyUj9ymKjutfu+pHFaTKU4lJD3f3L8rs/yZuCknLB2T2zv51qz5eeSIqzY8/YJ93u9m/JZ6CtrrsB5Lva9LTy2Wlvik0clHrJQPhiXfZg/Hl5jo6vK4Z6Gn53FRltXcNhZvcexYyGQj1CU+WO4u+fnxgo+smSf/c39v8/Ua6xmvbNJEM6M5CcfT59eLsuW/e+LeFYvErkoVu7NLCRuUEMKx5Ebr7XzewrqU0piLx9PrsxA3vBisU39C39rCh5PnD07NU+FYpuXHd5PGS4yZLJBu5rm1nb9jfdkSXy/RcSTcmezrGelZWXCDZGxspK3zHXF/xSGGdj9fO5DpPNTAxETfhFFP0wrm1dQ8eZG5PK9qSX4wuizdnIT4bteV76QVIBtIbLvoSyk9hoRv9rPLiMg/I/SJT4QPzOeq/9P4yGr2qtOiR6/IVW+5/ReWJr+6PvBSb61cuXSH5/SY72U1PKj3NlOe1P/tNv3nI0qXLL5DZK8tz3m9X7Geu+EvLVZUn2rGbTRy6a1dLNrC86cETp+6udjv69OEnxOYvl+wYWThOYbGdpvBO3eQ/cuniFZu3XVPvdUe9ZGD81YGrON0gzqPPuOOAOI6/Jfkz5H5af2/rhQsbVCFxG36WpP+qLFAVFheS/JwDtrR8YyY9BYlkQPVmsiR6I8LT/RUWFc+o3NQ7Oode76Yww+HHlRtvbZyEVxpT9bJ0dA52uJXGaOzlprcOrGv5RBrL6vJ8zcC61lC/8setY83Gw7zgR8ssrD64l8yOlnRgiDuLdI6Kdlv4chztLK7c1Lpp84rhkPTOKOlLGph0cV9Py+srNWjvHHy5wgwe6a8ht+vvaflZI+3NVnw20iK+k1aAZCCtYIaOH/uhLr5XsjeUn6iLcr/Eo8K7Zzdoze3IM/+8YunIyOHhna+ZBmP57eFHr3IDae8a/JhkYcT8tkh2fKXbvk4ysMtgrZNP9sKWlcPvclfoPv+r5KdItkbSi8vskw5QbD998AWK7GvhNcXkT+aVufPxf5WfUqe+mqWZAHp7X2/zxcmn5nrJQOL98fiNrV7hYZVIN33JTX9r0v/GFncszKucXWtT8l4x/DKXPjjx6sWGZf6uFZvu/1Ij4z4ST7Gh27zq5pR40n6Om61buem+tyXLnJgSq8dWYqUmTkbN9PK+dS1hfEXdT3kxpQtLXfGmn/rS+HlTvZ4JvUx3HbTPIYWdYZVOKxQjvzV2+/P3elrCzVWJOj11uvKSqzDW3rTHb97lHqDJ/l6vUUkDN13z4LLoxMriYRPJgA3PZBpxFuMzQz+1mawqyUAmL1u6SrefMfQkxf5RyY4tP10X5fqWR9G/NJIUhJ6GOGp6j5tCF+l+dWpTNOkXsflauf2NSWEOt0z6dF9P85pwI93liU/20YGeVe+o7ZoeGzC27yUynVhznikHJiYHyEnq27e45ITkvPHSO+FodK2bQmIUVtUrPcBJ2ihpIOzXINMOWXSQPPQY6BnlBCqW/HMrNm/rrNzI6iUDiSe2qu7lShvGRn2Pvj/xznxBXmk0EjljNjtPcfP3ja9KabpLsf/7ii3bLp4qKUgkA4cqspf1r20OPTSlT1v30FnmOi+siFmvt6YmJkrrH9T8tw07i/7s765v/dNU7UgMnl06WfIQxqjEHp83xYC7B9zsKnNd5PJ3mvT3Yz1UkycjVd35Ut0xAZW4qL2pT9WeGoOqeJqIPbtnJslAluOzkRjmOzMXIBmYudkecsSuXcOSQlJwYSHe+fbJxhS0dw4fJ/PPj3WFjn/CgLO7SoPQFAaf+eMnef8/VLSmI29Yd8jGxDv1MLBpS+3AqCRyzc02/GmDFe15feubS12wk306uodOcFdYH/6B5JiFXd/32rC7f0DL4y9P9hRZunk37fyQ3E8NyYNLvQM9zd0heambDHQPvUGuME7gRjNb6/K3yu0WyW8z6WSXnpwY73B7FPvrrrug9QeLKbjqvF4K1btdbt39vauunmrDorFXNPr3/p6W/wwHjQ2C04Bkh5j7i/t6W/vqtbW9ayjMpHilS9+8r7jp5Qcue+he5RkZYcxA3cSqtpxSD8RIdI1cRyQT0PC9ScabhD+F8TRbJL9TsoMlrUokiROncP9if2/La+u1vZFk4OjTB58ZR3Zl6Omq/FuY7prXlFtlMLHqoS+byVLY7XtAfE7nxt9nJkAyMDOvRfvtsIJY7E3PGFjX/OWZVLJ+17A2lrvWqxZcScxJDqOSw4/nxcVC9KmD7jr018n3yuWn+ZfIFG4EyZH+4yOZQx0rP/yS/XCywWbhe+Hm/ZDCyq+adMJY2/y8/p7Wt03XzqrNcxLLKydeIYSpjzNYC6FqUOT2yk2tKhkoL+KTfPKS+4UyCzMrKj0Qoeph5Pxtbnpfo13w07V3qr+H67xpxcYToqJunS6Jqi2nzuulWLLPjjz4wFtq15CoeYodfzqujFwPvS61vTRViV85Gagsd7wl3vrXxLVvKBkYi63y66mqZZOrrl8k2bDcP7Yz9ktrexvKPTfh9VZlnM1YNadYhrmRZCAxM+DQRnc2rCq35vyJ1xh/02h5JZ+JsQY3Lob4TBPbHDs3AiQDc+O4W0upDNRz0x/D09TP1h8xOtMK1ekafsDdTqvMeU8+bUl+ZVMhOm26eeml+fPRyEdk1pm4EY4vf9vRNfSFRqdZJZbPVePrvpdmU1zippMq72hrXh/cUzvLYTq3qtcW5afE0hbRY1vXhilmpRtg4p3sHebRkW7xwZI/rnQ/MQ0uK4z+cro5+NPVZSZ/b+sePM3cPhGZnzDV3glTlbnL6yXXdb48PinZm1KdDEwkeYlrPeXyxxMJov5UjKOjbrhg1e8bTRqrkoozhk9U7GHMyPhMluTaG3J//4ot2z403TiIjjXDj1LBL3XpqeXyJ42ZRpKBmsR2fX9Py+nTXcfy7IvrJAtrMdSMDZiI8UaT5HC+xRaf0xnw9/kXIBmYf+N5P0N7+YfPpFuWjETHXPWZVX+Z7UnLe9ufV+4O/2tkdnzYrnb8x8N160ymbI0tTPSQj5l7dzkhGH9aTEyLmnbOddUP4gw2URo/R3mRl1ijDyv4zrD2QIukGScD5aeq0rr+Jv08eI8s12Nqk4HEQkR7zzThmO21m+y4iZX29Nz0sxXGNywKPR2rJP/svcXNp08koMmb09gNfZ+lO+4a7+qfYmvfmhvl+LUZX4PBdcd0MzMqBuVdO68JKz9WFnJq7xoKa1Ks8UnGp0zmF7ri40LTZeMJwSRtSCYDYRphX2/zK+q9TphIbP13O4tqm3oMhFtb18YPmzysoRBSyV0GCo6vNdDAgMlKGxdTfM51vFPe7ARIBmbntqiOSmT5Ve/GZ1vJqhHU5afftq7hN5h0kWaxln15hb3wlHZccuBUIhmYtvs3uUpfvQGBk94IO4dfZeZhq93SOXbEy5+QuHFXvbZo1Kt2BPdO194Ta9KPvcJI9kCY9KG+npawJsJu+VSNz5jinfdMKpfoqdkloUr0AmwPuzgWl/rPK+/wp+rVqXqtEypTvumOr9onlcprpGej5in9HD+o+QP25+GvyNQexf78mY7PaD/99qcoKlxdGiszSTI61bv9pO34ksSmQ939XwZ6Wz82mX3Nq7nwtV0MEv/+dzbqs5jicyZxx3fnT4BkYP5sF6zksa7ZZZdL3ib3L67Ysu3U6bo/p6tcort27CY6uuR5Y/OZG3tfX1t+4r36LytzyCvJQOUJe7oejUTyMOWAw+S5EzetUjtGRpY8qTyAa2wWhPtJxWjJj5cUdz7+gb2j71WmbE3l09Y5+DYz+2gYyBhGty+JCpGbf1/y5uRiMok56jvSrog33fWa7u9t3UMnmYc9JLRDrhf297bcON0xU/29apxEzZNy1UI4rrOWLxm5qNIzYK5P9PU2v6XeE3PbWOIWBqeG5XzHk4GqZX4b7BWqNze/HNMvbvSGuUsMT7Mw1FSj/qvLqnrav1tx8dj+Cw67pfZ8NStEhlUMwz4Dhdp/g8/t/uOqid6uxv99Lqb4TBOLHDs3AiQDc+O420up3nAk7e53VV29pZvotpHlB5WffpdaZC/rW9t8w8wa7dbeNfRRc3v4PfGmNaFbeeKmUd312dE1dHIs/3dJh5ii29yLVw30rj63ZpfE8ZHqU960Jn7A6/UMhBvOaSpEW8rvl+8312fjeMn76q1cVx4D8TaNrXy4d6WXY6/teni9ZKBm1sLGKPZXzvSJdGbGk397bOvhfUPXfphK+Ye0mxMlRsXvV/u0X76ph96Y8cGeid0B606hLJcXBr+GUfzlQ8e2AU7u/TBV93uy9fUGMlYSQ5e+W4gKL53p9stt3RsOlkeXyG1dvf0jGk8GpNIsgELhmrC+RVhfwlV8aX/PYb+ttCH5aqKyd0Rxpx89tiiTb0i+Xqh6vTKDVwWLKT7nKs4pZ/YCJAOzt1t0R7adPnyERaVBU4eE1d3corMbWTegtiFtazY83gpR2GBn9UQXd9XTTJhK+D9RrA9ed0Hz/062bG34cRzduXRF0T0s2BJG7sut8JeVm+7tCz0XiSfI8dXT2ro2vCRS4RSXh53jkp/1+xaXvKWyoU+jy7nW9nBs8+UthWKxPBirVPw5heLOiyo/zOUTjq97X1prwO0ppXUGrLT+/d6V7yj2l/Vf0Pqdmu7tqgFyNUvahqmbA7LovSs2H/q9mazsNxfBVjN+Y6vc3z7dugH1zluagXLA0Hs8LCHtuqN2p8WqRKE8oLK8VHCYTvjIykyUsNOlLNpf8hfKdVydDXZKlsnXHI32ItVLBqpWxBybBvt+LzZ9fvIli8cW1SoUdxxibk+sWBRd/1t/++Fdk+ipBol2dA+/2t0/W56G+xdz73Gzr5dWM3R/b2mNhzDWZXl8YhikmdxYyE2vrCzZHOqVmK2xy9LfU8XOYorPuYhxypi9AMnA7O0W5ZE1a82HhX5uc3lYsa2vUCzeumRZvLXeD1T4wV3yYLG1ye0Ul7rKiwlVPUFOslBPmGJ4l8lvcdm9Mj1RsQ6WKawzX0oAkp/wYx4XlxwbfoDLXdfhaXD8B6yta/A7Jju+Hu6O5dHee40Un+Bu7/WinTNwQfNPp7oIySemyhPlkWuGDm4qWOgiD/Pgw6c0eDGRSE08mU5e+ANmOrtvXfOnd1lAqc5ywnWWUA4lPyjpNpnCnvZbFOs2RfYUuTeZfD9XKQEJyyEPu8evG+hdPTg3AbfrfhUuXWdmV7qKYWOfPydXkJw459iNcUlx5Bmx610mPa30tzqbUCWmzwXj8cGhHZ2DL3WzcPMLsy7qfcKGVCHZqiRc44lV4tXM+CyDqTzqLWAUvl/a/U+Fr7v0hMTx4bx/lukWuZpM+js37ScvxXByOmg4JCyZ/cm+3pbSQlq1n8SKm9OOgwmpcVvn0GvMLOwNUWnzeJH1dpOcKN++dm/x7ldVBm5Wj0OY2c6Xiys+5ybKKWXmAiQDMzfLwBFuHWcMPdWL9uFptzSdvDU/tqK9pnZOemm++sqNrzf5B+U6qEGMsBDR9+T65Iot2745sSPdhoMjj/pcviIsmBIV4zdVVit0d5mF8Ay7tZbD1HVTf2/L2E2ogU/VDaH8rrlmsaNQ/PhWsmODFEdeYm4vdSvd7MLNOCyrGz73ybXBI//86I4dlyTn1yeSjudPNtd7bJ7/8KtN9v7y2ILpWjDtAlDTFTDV30s3gOV7haTvrTO4jskiH5D83BWbW86r7eGomr7putwPbj5p4FzbGQ4++ow7HhvHxfMkhe2BwzvwUm+JFfwdfWtbft7WOfTW8niM8PWJZKC7EitqbeSd/1RP0ZPsmDgd573m+ppZdP51PYfeOllv2ESCa3c0uiJgWCPE40LoCQgrXYakYIvJP75Pcel5yVUzQwUTY2BqkiK3jq7h9S79Y6O9ZskGL7b4nO5i8Pe5FyAZmHvTRVVi6AbUDnth2NjHpCOn2GY11Hurm/WH+ejTdWOHH4/NB95xhC6CB8UAABcySURBVIrxyxSpo/Sud+JJauypV36lmb66fVnh55MNzAs356U71Gbm/2TuL60kAS6/IpLucVnVNq/9PS0Nx2zipvS4QnHnC65d/6jh0g9qeYpZ+UI1tMTtdBe1PIL+yOlWlSu5HTD8DJP+0c06JA+vdMJNcWz3vrBRU1ikaLl/e6o19aerT6N/D/W566A7nlgoFk8NPTIuhZ0EQ33qfcIy079z2cVebFo/1Y6A5YWdLjXZm/t6mkNvQNUnXPcw1mLZkh33JXuqal4xVM3DH+u9if+fXP/e39saVpac9FPa/MmLV8ptiVvx+fX2eyg/ER8rsxNMepar1BNQ2YZ4q1y/kPTVqFC48uF3H3JbI691JlZajJY3mgw0eq1KsVtaSlxhMbCwF8hLk/s0tHVvONw8+pIrPnWgZ/UPZ1Ju5buLLT5n0waOmZ1Awz+ssyueoxabQHhv/OeH7XNAGAGfrFvtj/JC1bu9c/A8mb2lcr7SZj0edw/0rh448sw/rWzaOXJ3si4zSQYma0PVAkbyTTNZxnWyMsOP6D0tdy2/+qMHb1sou/k6z3O6hh621KOqbuvKrnuN3BDT1KuRhXvSlL8Qx7Z1b3y+FB+1V2HkP+djYalwfZosXlb/dc5CtJBz7IkCJAN74lXNQJvCjAGZvcDH1vovfdx17kBvS9jetfQpvQf1+PfJ5sSxP+v6lGv41zx9Njx3PQOsma/iZMsZZ75hNACBRS5AMrDIL9CeVr2wZW8cKUzNG9+TPSwIFNaAr7cXe3vXUBg0MP5xj98TphmmcakZ4FaaXtjf23phmjI5dq4EqkbkT7sy5VydlXIQyLsAyUDeI2CB2n/0WbcfFI8U/m95kFSYXhY+N7vHl091c2/vGvpjeTra2BFup/X3Nqe6ce8yiHAWqyouEFsuT5MYkU8ykMsIoNG7Q4BkYHeo5+icbZ0bWt2iN0bSuyeabRtN/rZ4WXzNdIPk2ruGwvTB5AyCI+r1IMyUdHy9+zBXYYp15GdaLt9PLzC+MiXXJT0mJSDQoADJQINQfG1mAiEJkEVnmRTGBFRGaIfpghf1rWt+Y6OltXcNbZDUWv7+z/p7WsLCP6k/VUvcTrOFcuqTUcCMBJLb61aWrp5RAXwZAQRmLEAyMGMyDphKYLIkQNLnvKnwzoFPHnrHTATbuwbvcteBFllYE+BX/T0tfzOT4yf7bmk52KjpRpkODavoNbob3lycmzKmFiAZIEIQWHgBkoGFN98jzxjWM4h2ROe7lXoCEh+7xb14dpgqOJuGt3cNbSktyxo+Zr/rX9f8uNmUU3tMzbiBe2ezk91c1IMydhVIzPbYEjaCmnqLXwQRQGAuBEgG5kIxx2WEnoBI0Tm7JgG6x6Rz42XxRdONC5iKr71reFjysSWCXV/o720JG+3MySe5u15lz/s5KTjHhYT1Fu7d746HXvWZQ/862Sp90/HUbhHd17sqDCLlgwAC8yhAMjCPuHty0W2dG9ois7NcVruhULhrXx/WD5iL9fSrpxbaQH9Pc/tcuVbtgsiMgtSsiV3wXqYoOq5/7aofzabQ8WTA9Jsl8ehRV/c+umrhqdmUyTEIIDC1AMlADiIk3LhDM2fbVV8hKo0HiKITzRU2aakM6ksI2i0unTvQ03zZXLHOZzJQ2rFtJLpGriOYUZD+ik2s3+D7plnVMbEFcgOb/aSvNyUggMD4DjBQ7KkC7V3D/ZK3lS/1jJ+s2zpve5Gs6YBI/pL6vQAluSGXnT2XSUDlesxnMhDO0dE19EGX/lXMKEj9TyCxeuBhxTg66oYLVlWtHtnoCSrTPknQGhXjewikF6BnIL3hoi0hLPnr0iVVFXS/dOloofOqz6z6S1gNsGjxgZW/mxX2cy/+Xfj/zQpHjSURU3zCjnSKP562x2GqU9RMLTy9v6clbPc6Z5/y5i7XyrWdGQXpWRObQH1+xeb7/6myQ2WjJSe3nZb8vP6e1rc1eizfQwCB2QuQDMzebtEfGV4PmEX9yYqWdwUM28n+RtIspunZLSa/KO3AwEbxJno2wqrBcftcJx6J99zHK9Ix/WtbftVo3fjergJHdw//fex+haQDTPqWRYVTr1t7aJgR0tDnud1/XFXwnd8Nu2C66ZUD61q+0tCBfAkBBFIJkAyk4lv8B7d3D/9Q7v+QqqZud5j8q/FYL8BgqrJmeHByBcK52Jeg3unDTo6bVq5cOrDuwPtnWD2+Xkego3v41e4eti1eJukvMn1oZMf2nhsvfPx904G1dQ2dYlJYbvquNK8apjsPf1/8Asef+ftlV3zyMQ8++x2bHnLjh1dOGzuLv0WLu4YkA4v7+sxJ7Tq6h3/p7k+aWWF+rxTd5PLPDPS0XDyzY+fu28nXBO7xGwd6V180d6XPf0nhhyycJV8/Zm5t3cMvN9enJT20rPygpMsiRR+4rufQW+tNOzyma+iwonSVpEe76ZoHl0Un/uD8Vdvn/ypxhsUm0N41+GnJTgvziStD2+brYWCxtX131YdkYHfJL/B527sGXyPZW2vW+a+txaBJ34s9vq24ZPna737yEZsWuJq7nC5ryUB71/CLJX+Hy54SKV7q0lJZVNqfOW8/Zh1rhh8VN/kXzBV6pqLExd1i8mvc9RWP/EYd2LpZf9p4eGR+oZv+VtKou71hoLf5S7s7/jj/wgq0rfntCkVL/8csel69M/f3tHDPmqdLAuw8wS7WYsP0wNDVf+SZf1q5dGRk5XUXtPx6sdY11Kutc8NNZtFTw/9dLBbffMP6w9Yu5vp2dA5udbPK03BppaTSk01IBkrTO1tz9m/Orb1z+FluOt/GNpxKJgWTXcpZDT5czHFB3SYEjj/Tlz24MyTNelosdZj8QHdbaqZ9JZV60ib7jDxk7/3y1cu2cJGTsx+mhYPlTHMj0NE1/EWXvyqUZtKH+npa/m1uSp67Uto6N5xqFr3A3Y8xs4dVlzzRzVn675HOipqKX77uE4fdNXc1yEZJIRE1i8ImVaeEAYKTJAbf8GXxKWlWrcyGxp5Vy/IiZHvF7tvNoldLOsQ9/qlZ4RlS/DiZbfXY9zezR5THkswCwNf297S+eRYHckgDAiQDDSDxld0n0N493C33tbLS0/UOV/zqgZ7Vc7ao0Uxb1nbqhuVaqlYVon8w6VmS1oyXEaoYe9iZMWQu2+Xaa7LyzewXfeuanzLT8+8p3w9rEmwvNh1mip5cenUV+2aZfae/p/mm2S5jvKfYLNZ2HPv6W/YZ2We/lZFFfx97/CRT9GiZnh0WICvPUppd1cNdKOTM5c94WWP/fb1J18Yeb5rrmUSzq+yeexTJwJ57bfeIloXXGU07RzZWniYssvf1rW1+10I3rqNz40tjxU8207lTndtlm+XFtQO9q89t6x56jrnCNLmpPpe66/MDvS3fWug2cT4EKgKl1UXHVhUN05HvHPvvFp7wwx25st5IGMxZN8Gt6f+qD1tz0y+V7OXkOXGEmX4Qx/HVUvRT/l0sXIySDCycNWeapUBH1/AWl4/tXCgN9ve0rJ5lUTM6rLxo08mSwv9M8vFNMvusXLdJ+ll/T8vPkl9MrpMwxckv7e9pecWMKseXEZilQHl58rbywmItkh6QNMPZRqWT3ynZ78oJw/4m/42bDbv7seZ2kEy3u/th5VcD9W78gya7zGP/Vv8FLdeFNSaieKSw0NOXZ8m4xx1GMrDHXdI9r0HtXUNhFcXkDfmI2pvuXLV6bF18nT7p6ovuD7j7NTJdVvDox40MwOzo3HiIK/5nNz0h7MBobo+Q6aCaOv/K5G/v62kNC/bwQWDOBMKYlsiibS7rrh/X9iMpPkyKSgtuuRevD70D7vEjqyrh2i5T2Hxq0Gzpsv6eQ347m0py05+N2vwfQzIw/8acIaVAW+fQi0z6ZhhBWP5sc9dHBnpb3pOy6NLhHV2Dx8uiV7h7mNYWRrzv8glPMLEXL1822vSNsJRzmvOWB9J9uF6Pg0k/ltmvYxU/oVj7L8b3pGE0+L17bV66ZOT+AgP90kTC3B9beeoPJe+6pLhvkutii6KHufQTj4u/Xrrt3p9c/YWnbJv7mlBi1gRIBrJ2xXJY3/JueOX3mGMA4b1i37qWMIBvxp/yngyvCD+W7nFbacBf/c+gTJd4HPfMR9dl+TXExyWVnsAme+9qZqWFltx9RLLHmqlqFcg4Lg5FZj+KZaOFWHdW9puYi0QiOUp8DD46z+RPHd/jzEweF98TxkjM+EJwQCqB0q6bo4WjFMcnSFZUadaN7TdRqA24Fz8XnuTnIhZSVZaDF70AycCiv0RUMAgc1Tl4XmT2lioNi17ev27V1xoVKu3VEBVeL/c3VY6pcwO+yj3+YcGjSxp5BdDouSf73lhCMNZ929AgrDoFTTeSOyQTIWEoH1pKJEyFKHL/YUgcKk+HbWfecahGRx9tip7uZm0mP37qsss1JiFIGwZTHh9u+hoJg/uio+QKC1WEAX37j28jbvZTuW9wj8OaIQPc+Of1cuyxhZMM7LGXds9rWGUw3sTUI9vucfHDsuh7Jv3VPQ4b4pRGRIf/HUWFTXFc3GYWhXnt4b/X/Zjp9jiOv6Co6eKBdav+sDvkwu6J8ugUxf4Qi+zoqepbW7/ZJhHj5bjvkFl4X7zLK5Lpy04sF6v4QwM9qxfdOhC743rO9pzJG3/kIUG0wxOxcE+53MvcbEBxcVDLdTOvamarzXFJAZIB4iEzAh3dG17tHv1PIxWe/iZW6nc/3+XfWIxPUuWpXqeGtkZRoSV2D7sA/qnymsDd93GPvxMGhsUev7AQWTGOtcOiaG/3kPhMs/309Ig/k+w+yfdXpBvk2uFxvN2iQlj3YeWkh7t2yPQbc90cR7pZcXzLYvSdvvkL841SEqjoqCjW4W4KN/7wP4mPX29mA7HbzfLizfPxumphWspZFrsAycBiv0LUr0qgvWvodknTTi2cIhn4lbsunavBh4v98pQHlA2aRQe4h3fLE+MNLCrs7XHxwNC1bBbdJ7enuIp9oU1T3XQ6zhx+lI36MbH0AZkqUz6noQjvr31AUXz5wLrVNy92t/moXynBs3JXv1mb5OHGH7r7K58huW4200DscbjxD8xHPSgTgXoCJAPEReYE2rqHrjZX3Y1MqhtjGyW/Qma3hJveQo0DyBzoLCtcGoNhUf8sDt9qrsu8YJf1r22+fBbHL/pDwhO/WaHFix42YDrcFW7+VTf+0Iah4BBHdrPiYnjXv6Dbgy96RCq4oAIkAwvKzcnmUiAMvosVHyIVDjH5fma2NHSRh3nSDKSaS+nJy2pwUaWGeg3MooH+nlXh2mXmU3XTD8vyjnX1h/EpySf+cnv8enO7OQ7v++nyz8w1zktFSQbycqVpJwLzJFBKyjze5OYrzZr2ieRhrfo2yY6azSlNflksC0/JNxdV/Pl3ew77xWzKmYtjSl37kfYff8oPN3rz/WsG9tU5VfnGP/bUz9S+ubgYlDGvAiQD88pL4QjkWyA8OUcetbmrTVaa5ZHY3nnGNoOSDZrirXLbGofxD5G2yrXVItuq2LaGEt13Dk3W5T4+Wj8ON/gmd8WtMu0fFngK/9vGRu8vl3z5roP5autrt0jx1tLTfvmmz+j+GV9TDlgkAiQDi+RCUA0E8iAwNnq+cLiVeg68dba9B/Np5dJ2k+6WQnIxdrOXaTC2kHAUQ49FeNLn/f58XgTKXnABkoEFJ+eECCCQFCg9re9ILKKT+KMv8aKNWiGy6PDKE7yXu+rHvlZ6kp9Nb8M9kpdmNZSm7oXXEh56GYpbGcVPfOZRgGQgj1edNiOwBwqUeh1Cd/9kn/BKYakGWaRnD7z4NCm1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUguQDKQmpAAEEEAAAQSyLUAykO3rR+0RQAABBBBILUAykJqQAhBAAAEEEMi2AMlAtq8ftUcAAQQQQCC1AMlAakIKQAABBBBAINsCJAPZvn7UHgEEEEAAgdQCJAOpCSkAAQQQQACBbAuQDGT7+lF7BBBAAAEEUgv8fwA4c3jZPFf8AAAAAElFTkSuQmCC",
"x": 10,
"y": 230,
"pages": "0-"
}
]
}
```
## `Example` Response (B)
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/03c5c55183c74f8d94a4ec952e4e32ad/f1040-form-filled.pdf",
"pageCount": 3,
"error": false,
"status": 200,
"name": "f1040-form-filled",
"remainingCredits": 60822
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"async": false,
"inline": true,
"name": "newDocument",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf",
"annotations": [
{
"text": "Sample Text 1",
"x": 150,
"y": 100,
"size": 20,
"pages": "0-"
},
{
"text": "sample text that is centered (can also set right or left alignment) ",
"x": "10",
"y": "10",
"width": "500",
"height": "200",
"size": "7",
"pages": "0",
"alignment": "center"
},
{
"text": "Sample Text 2 - Click here to test link\r\n(CLICK ME!)",
"x": 250,
"y": 240,
"size": 24,
"pages": "0-",
"color": "CCBBAA",
"link": "https://pdf.co/",
"fontName": "Comic Sans MS",
"fontItalic": true,
"fontBold": true,
"fontStrikeout": false,
"fontUnderline": true
},
{
"text": "Simple text 3",
"x": 100,
"y": 230,
"size": 12,
"pages": "0-",
"type": "Text"
},
{
"text": "sample text 3 - input text field",
"x": 100,
"y": 170,
"size": 16,
"pages": "0-",
"type": "TextField",
"id": "textfield1"
},
{
"x": 200,
"y": 120,
"size": 16,
"pages": "0-",
"type": "Checkbox",
"id": "checkbox2"
},
{
"x": 200,
"y": 140,
"size": 16,
"pages": "0-",
"type": "CheckboxChecked",
"id": "checkbox3"
}
```
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"async": false,
"inline": true,
"name": "f1040-form-filled",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true",
"imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200",
"fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Text annotation params
const Type = "annotation";
const X = 400;
const Y = 600;
const Text = "APPROVED";
const FontName = "Times New Roman";
const FontSize = 24;
const Color = "FF0000";
// * Add Text *
// Prepare request to `PDF Edit` API endpoint
var queryPath = `/v1/pdf/edit/add`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile),
url: SourceFileUrl,
password: Password,
annotations:[
{
pages: Pages,
x: X,
y: Y,
text: Text,
fontname: FontName,
size: FontSize,
color: Color
}
]
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: encodeURI(queryPath),
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download the output file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file).on("close", () => {
console.log(`Generated PDF file saved to '${DestinationFile}' file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.error(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
'https://api.pdf.co/v1/pdf/edit/add',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => '',
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_POSTFIELDS =>'{
"async": false,
"encrypt": false,
"inline": true,
"name": "f1040-form-filled",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true",
"imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200",
"fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789"
}',
CURLOPT_HTTPHEADER => array(
'Content-Type: application/json',
'x-api-key: '
),
));
$response = curl_exec($curl);
curl_close($curl);
echo $response;
?>
```
```csharp theme={null}
using Newtonsoft.Json.Linq;
using System;
using System.Globalization;
using System.IO;
using System.Net;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*****************************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
// Text annotation params
const string Type = "annotation";
const int X = 400;
const int Y = 600;
const string Text = "APPROVED";
const string FontName = "Times New Roman";
const float FontSize = 24;
const string FontColor = "FF0000";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// * Add text annotation *
// Prepare requests params as JSON
// See documentation: https://developer.pdf.co
string jsonPayload = $@"{{
""name"": ""{Path.GetFileName(DestinationFile)}"",
""url"": ""{SourceFileUrl}"",
""password"": ""{Password}"",
""annotations"": [
{{
""x"": {X},
""y"": {Y},
""text"": ""{Text}"",
""fontname"": ""{FontName}"",
""size"": ""{FontSize.ToString(CultureInfo.InvariantCulture)}"",
""color"": ""{FontColor}"",
""pages"": ""{Pages}""
}}
]
}}"; ;
try
{
// URL of "PDF Edit" endpoint
string url = "https://api.pdf.co/v1/pdf/edit/add";
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download generated PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
finally
{
webClient.Dispose();
}
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination PDF file name
final static Path ResultFile = Paths.get(".\\result.pdf");
// Text annotation params
private final static String Type2 = "annotation";
private final static int X2 = 400;
private final static int Y2 = 600;
private final static String Text = "APPROVED";
private final static String FontName = "Times New Roman";
private final static float FontSize = 24;
private final static String Color = "FF0000";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// * Add text annotation *
// Prepare URL for `PDF Edit` API call
String query = "https://api.pdf.co/v1/pdf/edit/add";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{
\"name\": \"%s\",
\"url\": \"%s\",
\"password\": \"%s\",
annotations:[{
\"pages\": \"%s\",
\"text\": \"%s\",
\"x\": \"%s\",
\"y\": \"%s\",
\"fontname\": \"%s\",
\"size\": \"%s\",
\"color\": \"%s\"
}]
}",
ResultFile.getFileName(),
SourceFileUrl,
Password,
Pages,
Text,
X2,
Y2,
FontName,
FontSize,
Color);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated output file
String resultFileUrl = json.get("url").getAsString();
// Download the image file
downloadFile(webClient, resultFileUrl, ResultFile);
System.out.printf("Generated file saved to \"%s\" file.", ResultFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile.toFile());
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
'https://api.pdf.co/v1/pdf/edit/add',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => '',
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_POSTFIELDS =>'{
"async": false,
"encrypt": false,
"inline": true,
"name": "f1040-form-filled",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"annotationsString": "250;20;0-;PDF form filled with PDF.co API;24+bold+italic+underline+strikeout;Arial;FF0000;www.pdf.co;true",
"imagesString": "100;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png|400;180;0-;pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png;www.pdf.co;200;200",
"fieldsString": "1;topmostSubform[0].Page1[0].f1_02[0];John A. Doe|1;topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1];true|1;topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0];123456789"
}',
CURLOPT_HTTPHEADER => array(
'Content-Type: application/json',
'x-api-key: '
),
));
$response = curl_exec($curl);
curl_close($curl);
echo $response;
?>
```
***
## Create Fillable PDF Forms
You can create fillable PDF forms by adding editable text boxes and checkboxes.
By using the [annotations\[\]](/api/pdf-add#pdf-add-annotations) attribute and setting the type to textfield or checkbox you can create form elements to be placed on your PDF .
### `Example` Payload (Create Fillable PDF Forms)
```json theme={null}
{
"async": false,
"inline": true,
"name": "newDocument",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf",
"annotations":[
{
"text":"sample prefilled text",
"x": 10,
"y": 30,
"size": 12,
"pages": "0-",
"type": "TextField",
"id": "textfield1"
},
{
"x": 100,
"y": 150,
"size": 12,
"pages": "0-",
"type": "Checkbox",
"id": "checkbox2"
},
{
"x": 100,
"y": 170,
"size": 12,
"pages": "0-",
"link": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png",
"type": "CheckboxChecked",
"id":"checkbox3"
}
]
}
```
### `Example` Response (Create Fillable PDF Forms)
```json theme={null}
{
"url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/d5c6efa549194ffaacb2eedd318e0320/newDocument.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzECMaDJJV7qKrpnGUrZHrwSKBATR5rxVlQoU0zj3r4jyHPt7yj4HoCIBi65IbMRWVX8qZZtKL9YGUzP%2FcemlqVd4Vi5%2B80Sg%2BymqQtaQ8qSFqKA82JnV%2BNBDatIigZIZha%2BrQM3jSC%2FZhX1zxsfLLsaH3K5nBnkjT3gi%2FZnx%2FgqrlIhf3m2xRFaTlgHrBADlK9KKPIijSusD4BTIo%2FQ433xx%2FQEaGWdX0nu4NuiByyXNPsBCAI3im9LMUCujjqF79ocyLHA%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHCSWKUQ4T/20200716/us-west-2/s3/aws4_request&X-Amz-Date=20200716T092641Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=2aa88d39aaf4b5891e4cb42d5675a64486098558d7159b37b75252209bdd6a95",
"pageCount": 1,
"error": false,
"status": 200,
"name": "newDocument",
"remainingCredits": 77762
}
```
### CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"async": false,
"inline": true,
"name": "newDocument",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/sample.pdf",
"annotations":[
{
"text":"sample prefilled text",
"x": 10,
"y": 30,
"size": 12,
"pages": "0-",
"type": "TextField",
"id": "textfield1"
},
{
"x": 100,
"y": 150,
"size": 12,
"pages": "0-",
"type": "Checkbox",
"id": "checkbox2"
},
{
"x": 100,
"y": 170,
"size": 12,
"pages": "0-",
"link": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png",
"type": "CheckboxChecked",
"id":"checkbox3"
}
]
}'
```
## Fill PDF Forms
You can fill existing form fields in a PDF after identifying the form field names.
Once form fields are identified then the [fields\[\]](#:~:text=height%20parameters%20provided.-,fields,-array%5Bobject%5D) attribute should be used to populate the fields by fieldName .
### `Example` Payload (Fill PDF Forms)
```json theme={null}
{
"async": false,
"inline": true,
"name": "f1040-filled",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"fields": [
{
"fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]",
"pages": "1",
"text": "True"
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_02[0]",
"pages": "1",
"text": "John A."
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_03[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]",
"pages": "1",
"text": "123456789"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B.",
"fontName": "Arial",
"size": 6,
"fontBold": true,
"fontItalic": true,
"fontStrikeout": true,
"fontUnderline": true
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B."
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]",
"pages": "1",
"text": "987654321"
}
]
```
### `Example` Response (Fill PDF Forms)
```json theme={null}
{
"hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"url": "https://pdf-temp-files.s3.amazonaws.com/cd15a09771554bed88d6419c1e2f2b16/f1040-filled.pdf",
"pageCount": 3,
"error": false,
"status": 200,
"name": "f1040-filled.pdf",
"remainingCredits": 99999369,
"credits": 63
}
```
### CURL
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/add' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"async": false,
"inline": true,
"name": "f1040-filled",
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf",
"fields": [
{
"fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]",
"pages": "1",
"text": "True"
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_02[0]",
"pages": "1",
"text": "John A."
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_03[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]",
"pages": "1",
"text": "123456789"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B.",
"fontName": "Arial",
"size": 6,
"fontBold": true,
"fontItalic": true,
"fontStrikeout": true,
"fontUnderline": true
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B."
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]",
"pages": "1",
"text": "987654321"
}
],
"annotations":[
{
"text":"Sample Filled with PDF.co API using /pdf/edit/add. Get fields from forms using /pdf/info/fields. This text is be added on the first (0) and the last (!0) pages.",
"x": 400,
"y": 10,
"width": 200,
"height": 500,
"size": 12,
"pages": "0-",
"color": "FF0000",
"link": "https://pdf.co"
}
]
}'
```
### Code samples (For Fill PDF Forms)
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Runs processing asynchronously. Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success). Must be one of: true, false.
const async = false;
// Form field data
var fields = [
{
"fieldName": "topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]",
"pages": "1",
"text": "True"
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_02[0]",
"pages": "1",
"text": "John A."
},
{
"fieldName": "topmostSubform[0].Page1[0].f1_03[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]",
"pages": "1",
"text": "123456789"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B."
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]",
"pages": "1",
"text": "Joan B."
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]",
"pages": "1",
"text": "Doe"
},
{
"fieldName": "topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]",
"pages": "1",
"text": "987654321"
}
];
// * Fill forms *
// Prepare request to `PDF Edit` API endpoint
var queryPath = `/v1/pdf/edit/add`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile),
password: Password,
url: SourceFileUrl,
async: async,
fields: fields
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download the PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file).on("close", () => {
console.log(`Generated PDF file saved to '${DestinationFile}' file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.error(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "**************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
def main(args = None):
fillPDFForm()
def fillPDFForm():
"""Fill PDF form using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
payload = "{\n \"async\": false,\n \"encrypt\": false,\n \"name\": \"f1040-filled\",\n \"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf\",\n \"fields\": [\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]\",\n \"pages\": \"1\",\n \"text\": \"True\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].f1_02[0]\",\n \"pages\": \"1\",\n \"text\": \"John A.\"\n }, \n {\n \"fieldName\": \"topmostSubform[0].Page1[0].f1_03[0]\",\n \"pages\": \"1\",\n \"text\": \"Doe\"\n }, \n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]\",\n \"pages\": \"1\",\n \"text\": \"123456789\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n \"pages\": \"1\",\n \"text\": \"Joan B.\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n \"pages\": \"1\",\n \"text\": \"Joan B.\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]\",\n \"pages\": \"1\",\n \"text\": \"Doe\"\n },\n {\n \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]\",\n \"pages\": \"1\",\n \"text\": \"987654321\"\n } \n\n\n\n ],\n \"annotations\":[\n {\n \"text\":\"Sample Filled with PDF.co API using /pdf/edit/add. Get fields from forms using /pdf/info/fields\",\n \"x\": 10,\n \"y\": 10,\n \"size\": 12,\n \"pages\": \"0-\",\n \"color\": \"FFCCCC\",\n \"link\": \"https://pdf.co\"\n }\n ], \n \"images\": [ \n ]\n}"
# Prepare URL for 'Fill PDF' API request
url = "{}/pdf/edit/add".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=payload, headers={"x-api-key": API_KEY, 'Content-Type': 'application/json'})
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.Net;
using System.Runtime.InteropServices;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*************************";
// Direct URL of source PDF file.
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// File name for generated output. Must be a String
const string FileName = "f1040-form-filled";
// Destination File Name
const string DestinationFile = "./result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Values to fill out pdf fields with built-in pdf form filler
var fields = new List
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "****************************";
// Direct URL of source PDF file.
final static String SourceFileUrl = "pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-form/f1040.pdf";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination PDF file name
final static Path ResultFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `PDF Edit` API call
String query = "https://api.pdf.co/v1/pdf/edit/add";
// Prepare form filling data
String fields = "[\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].FilingStatus[0].c1_01[1]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"True\"\n" +
" },\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].f1_02[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"John A.\"\n" +
" }, \n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].f1_03[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"Doe\"\n" +
" }, \n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_04[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"123456789\"\n" +
" },\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"Joan B.\"\n" +
" },\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_05[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"Joan B.\"\n" +
" },\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_06[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"Doe\"\n" +
" },\n" +
" {\n" +
" \"fieldName\": \"topmostSubform[0].Page1[0].YourSocial_ReadOrderControl[0].f1_07[0]\",\n" +
" \"pages\": \"1\",\n" +
" \"text\": \"987654321\"\n" +
" } \n" +
" ]";
// Asynchronous Job
String async = "false";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\n" +
" \"url\": \"%s\",\n" +
" \"async\": %s,\n" +
" \"encrypt\": false,\n" +
" \"name\": \"f1040-filled\",\n" +
" \"fields\": %s"+
"}", SourceFileUrl, async, fields);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated output file
String resultFileUrl = json.get("url").getAsString();
// Download the image file
downloadFile(webClient, resultFileUrl, ResultFile);
System.out.printf("Generated file saved to \"%s\" file.", ResultFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, Path destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile.toFile());
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
```
# Make Text Searchable
Source: https://developer.pdf.co/api/pdf-change-text-searchable/searchable
Convert scanned PDF or image files into text-searchable PDFs by running OCR and adding an invisible text layer for search and indexing.
**Try it live:** [Make Text Searchable → API Tester](/api-tester/pdf-change-text-searchable/searchable) — send a real request from your browser.
## `POST /v1/pdf/makesearchable`
This endpoint uses **300 DPI rendering** by default, which may significantly increase output file size-especially for image-heavy PDFs. To reduce file size, set a lower rendering resolution by using the [`profiles`](/api/profiles) attribute in the request body. For example, to set the rendering resolution to 72 DPI, use:
```json theme={null}
{ 'RenderingResolution': 72 }
```
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ---------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-make-searchable/sample.pdf",
"lang": "eng",
"pages": "",
"name": "result.pdf",
"password": "",
"async": "false",
"profiles": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/a0d52f35504e47148d1771fce875db7b/result.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 99033681,
"credits": 35
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/makesearchable' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-make-searchable/sample.pdf",
"lang": "eng",
"pages": "",
"name": "result.pdf",
"password": "",
"async": "false",
"profiles": ""
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more.
const Language = "eng";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. MAKE UPLOADED PDF FILE SEARCHABLE
makePdfSearchable(API_KEY, uploadedFileUrl, Password, Pages, Language, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function makePdfSearchable(apiKey, uploadedFileUrl, password, pages, language, destinationFile) {
// Prepare request to `Make Searchable PDF` API endpoint
var queryPath = `/v1/pdf/makesearchable`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, lang: language, url: uploadedFileUrl, async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url, destinationFile);
}
else {
// Service reported error
console.log("makePdfSearchable(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("makePdfSearchable(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(destinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${destinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more.
Language = "eng"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
makeSearchablePDF(uploadedFileUrl, DestinationFile)
def makeSearchablePDF(uploadedFileUrl, destinationFile):
"""Make Uploaded PDF file Searchable using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["lang"] = Language
parameters["url"] = uploadedFileUrl
# Prepare URL for 'Make Searchable PDF' API request
url = "{}/pdf/makesearchable".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more.
const string Language = "eng";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. MAKE UPLOADED PDF FILE SEARCHABLE
// URL for `Make Searchable PDF` API call
var url = "https://api.pdf.co/v1/pdf/makesearchable";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("url", uploadedFileUrl);
parameters.Add("pages", Pages);
parameters.Add("lang", Language);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// OCR language. "eng", "fra", "deu", "spa" supported currently. Let us know if you need more.
final static String Language = "eng";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. MAKE UPLOADED PDF FILE SEARCHABLE
MakePdfSearchable(webClient, API_KEY, DestinationFile, Password, Pages, Language, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void MakePdfSearchable(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String language, String uploadedFileUrl) throws IOException
{
// Prepare URL for `Make Searchable PDF` API call
String query = "https://api.pdf.co/v1/pdf/makesearchable";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"lang\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
language,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Make PDF Searchable Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# Make Text Unsearchable
Source: https://developer.pdf.co/api/pdf-change-text-searchable/unsearchable
Convert a PDF into a flat, image-only version that is no longer text-searchable, by rasterizing each page as a scanned image.
**Try it live:** [Make Text Unsearchable → API Tester](/api-tester/pdf-change-text-searchable/unsearchable) — send a real request from your browser.
## `POST /v1/pdf/makeunsearchable`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ---------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"pages": "",
"name": "result.pdf",
"password": "",
"async": "false",
"profiles": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/6b755238963a472abf67fd5e7ffafd79/result.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 327244,
"credits": 35
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/makeunsearchable' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"pages": "",
"name": "result.pdf",
"password": "",
"async": "false",
"profiles": ""
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. MAKE UPLOADED PDF FILE UNSEARCHABLE
makePdfUnSearchable(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function makePdfUnSearchable(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `Make UnSearchable PDF` API endpoint
var queryPath = `/v1/pdf/makeunsearchable`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl, async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url, destinationFile);
}
else {
// Service reported error
console.log("makePdfUnSearchable(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("makePdfUnSearchable(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(destinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${destinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import requests
import json
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "*****************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
fileName = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf"
url = "{}/pdf/makeunsearchable?url={}".format(BASE_URL, fileName)
# Execute request and get response as JSON
response = requests.get(url, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL of unsearchable PDF
unsearchableFile = json["url"]
print(unsearchableFile)
```
```php theme={null}
$apiKey = "***************";
$fileName = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/document-parser/sample-invoice.pdf";
$url = "https://api.pdf.co/v1/pdf/makeunsearchable?url=" . $fileName);
// Create request
$curl = curl_init();
curl_setopt($curl, CURLOPT_HTTPHEADER, array("x-api-key: " . $apiKey));
curl_setopt($curl, CURLOPT_URL, $url);
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);
// Execute request
$result = curl_exec($curl);
?>
```
# PDF Compress
Source: https://developer.pdf.co/api/pdf-compress
POST /v2/pdf/compress
Compress a PDF with ready-to-use presets or advanced image and font controls.
**Try it live:** [PDF Compress → API Tester](/api-tester/pdf-compress) — send a real request from your browser.
## `POST /v2/pdf/compress`
This is the current PDF compression endpoint. The legacy PDF Optimize V1 endpoint is deprecated.
## Quick start
A request containing only `url` uses the standard configuration — the same image optimization as Adobe Acrobat Pro — so you can start compressing immediately.
Start with the `medium` preset for balanced compression. Choose another preset only when you need lighter or stronger compression.
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
## Choose a preset
For most files, choose one of four `compression_level` values. You do not need to understand the advanced configuration to use them.
* `low` — Light compression with the highest image resolution.
* `medium` — Balanced compression based on Adobe Standard resolution targets.
* `high` — Stronger compression with lower image resolution.
* `aggressive` — The smallest images and strongest compression of the four presets.
You can also set `color_quality` from `1` to `100`. Higher values preserve more color and grayscale image quality and usually produce larger files. The default is `80`.
```json Preset request body theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
"compression_level": "high",
"color_quality": 80
}
```
Exact PPI thresholds used by each preset.
| Preset | Color and grayscale | Monochrome |
| ------------ | -------------------------- | -------------------------- |
| `low` | 200 PPI when above 300 PPI | 400 PPI when above 600 PPI |
| `medium` | 150 PPI when above 225 PPI | 300 PPI when above 450 PPI |
| `high` | 127 PPI when above 172 PPI | 180 PPI when above 270 PPI |
| `aggressive` | 96 PPI when above 120 PPI | 100 PPI when above 150 PPI |
## Request body
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }` There are no query parameters.
URL of the source PDF. The endpoint processes the complete document. The URL must be reachable by PDF.co — see [supported file sources](/api/url-input-and-request-limits#supported-file-sources). For a protected source location, provide a temporary or presigned URL.
Ready-to-use compression profile: `low`, `medium`, `high`, or `aggressive`. When omitted, the standard configuration is applied. Presets also enable font subsetting and font stream compression.
JPEG2000 quality for color and grayscale images, from `1` for the smallest file and lowest quality to `100` for the highest quality. It can be used with or without `compression_level`.
Password for opening an encrypted source PDF. Omit it for an unprotected PDF.
Set to `true` for large or long-running documents. The initial response includes `jobId`; use the [Background Job Check endpoint](/api/job-check) to retrieve the final status. Also see [Webhooks & Callbacks](/api/webhooks).
Callback URL notified when an asynchronous job finishes. Use it with `async: true`.
Output file name. The endpoint appends `.pdf` when needed. If omitted, it derives the name from `url` when possible.
Number of minutes before the temporary output URL expires. After this period, generated files are deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum retention depends on your subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files).
The current V2 Compress implementation does not use `httpusername` or `httppassword`. The old `profiles.outputDataFormat` and `profiles.JPEGQuality` descriptions also do not apply to this endpoint.
## Advanced configuration
Use `config` only when a preset is not enough. It is a partial object: you only send the values you want to override.
1. PDF.co starts with the standard configuration.
2. It applies the selected `compression_level`, when supplied.
3. It applies `color_quality`, when supplied.
4. It deep-merges `config` as the final override.
| Request | Image resolution | Color and grayscale encoding | Fonts |
| ------------------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------- | ------------------- |
| `url` only | Standard targets | JPEG, quality 60 | Unchanged |
| `compression_level` | Targets from the selected preset | JPEG2000, quality 80 unless overridden | Subset and compress |
| `color_quality` only | Standard targets | JPEG2000 at the selected quality | Unchanged |
| `config` only | Start with the standard configuration, then replace only the values supplied in `config`. | | |
| Preset or quality plus `config` | Apply the preset and quality first, then replace only the values supplied in `config`. | | |
Specify the narrowest override you need. Every omitted value continues to come from the base configuration, or from the preset and `color_quality` you selected.
Each preset resolves to an effective configuration for the first compression pass. All presets use JPEG2000 quality 80 for color and grayscale images, CCITT Group 4 for monochrome images, font subsetting and compression, and garbage collection level 4. Their resolution targets differ per the **Preset resolution targets** table above.
To represent another preset, change the four PPI values using that table. A `color_quality` of 80 maps to JPEG2000 `rates` with `quality_layers: [20]`.
The effective values produced by the `medium` preset:
```json theme={null}
{
"config": {
"images": {
"color": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 150,
"threshold_ppi": 225
},
"compression": {
"skip": false,
"compression_format": "jpeg2000",
"compression_params": {
"quality_mode": "rates",
"quality_layers": [20]
}
}
},
"grayscale": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 150,
"threshold_ppi": 225
},
"compression": {
"skip": false,
"compression_format": "jpeg2000",
"compression_params": {
"quality_mode": "rates",
"quality_layers": [20]
}
}
},
"monochrome": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 300,
"threshold_ppi": 450
},
"compression": {
"skip": false,
"compression_format": "ccitt_g4",
"compression_params": {}
}
}
},
"fonts": {
"subset": true,
"compress": true
},
"save": {
"garbage": 4
}
}
}
```
Use `compression_level` when a preset already fits your needs. The configuration above produces the same first pass, but explicit `config` values are preserved during the standard fallback retry, so manually copying a complete preset can change fallback behavior.
### Common config recipes
These examples show only the request fields relevant to the change. Add the same fields to the request body together with `url`.
Keep the base encoding and fonts, but retain more color and grayscale detail. Images are reduced to 200 PPI only when their effective resolution is above 300 PPI. Monochrome images keep the base settings.
```json theme={null}
{
"config": {
"images": {
"color": {
"downsample": {
"downsample_ppi": 200,
"threshold_ppi": 300
}
},
"grayscale": {
"downsample": {
"downsample_ppi": 200,
"threshold_ppi": 300
}
}
}
}
}
```
Preserve pixel dimensions while still applying JPEG2000 compression.
```json theme={null}
{
"config": {
"images": {
"color": { "downsample": { "skip": true } },
"grayscale": { "downsample": { "skip": true } }
}
}
}
```
Do not downsample or re-encode monochrome images.
```json theme={null}
{
"config": {
"images": {
"monochrome": { "skip": true }
}
}
}
```
Start with `high`, preserve more image quality, and make two exceptions. This keeps the `high` resolution targets, uses color quality 90, leaves monochrome images unchanged, and disables font subsetting. Font stream compression remains enabled.
```json theme={null}
{
"compression_level": "high",
"color_quality": 90,
"config": {
"images": {
"monochrome": { "skip": true }
},
"fonts": {
"subset": false
}
}
}
```
### What each skip setting does
| Setting | What it skips | What can still run |
| -------------------- | ------------------------------------------------------ | ------------------------------------------------------------ |
| `images..skip` | All optimization for that image type | Other image types and fonts |
| `downsample.skip` | Resolution reduction | Image re-encoding |
| `compression.skip` | The selected JPEG, JPEG2000, CCITT, or ZIP compression | Downsampling; resized image data may still be written as PNG |
Advanced image, font, and save controls.
Settings for color, grayscale, and monochrome images.
`color`, `grayscale` & `monochrome` all use the same object schema:
Skip both downsampling and re-encoding for this image type.
Control resolution reduction.
Preserve image dimensions while allowing re-encoding.
Target resolution when the image is above `threshold_ppi`. Default is `150` for `color` and `grayscale`, `300` for `monochrome`.
Minimum effective resolution that triggers downsampling. Default is `225` for `color` and `grayscale`, `450` for `monochrome`.
Control image re-encoding.
Skip the selected compression format while allowing downsampling. Resized image data may still be written as PNG.
`jpeg`, `jpeg2000`, `ccitt_g4`, `ccitt_g3`, or `zip`. Use CCITT for monochrome images.
JPEG or JPEG2000 quality settings.
JPEG quality from `1` to `100`. This field is used only when `compression_format` is `jpeg`.
JPEG2000 mode: `rates` or `dB`.
JPEG2000 quality layers. When supplied directly without valid layers, the fallback is `[30]` for `rates` or `[38.0, 34.0, 30.0]` for `dB`.
Same object schema as `color`.
Same object schema as `color`.
Font subsetting and stream compression.
Remove unused glyphs from fonts that can be subset.
Compress font streams while saving the PDF.
PDF cleanup settings.
Garbage collection level from `0` to `4`.
* `0` — none
* `1` — remove unused objects
* `2` — compact xref
* `3` — merge duplicate objects
* `4` — detect duplicate stream content
Fallback base used when the first compression pass is not smaller. This is also the configuration applied when the request contains only `url`.
```json theme={null}
{
"images": {
"color": {
"skip": false,
"downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 },
"compression": { "skip": false, "compression_format": "jpeg", "compression_params": { "quality": 60 } }
},
"grayscale": {
"skip": false,
"downsample": { "skip": false, "downsample_ppi": 150, "threshold_ppi": 225 },
"compression": { "skip": false, "compression_format": "jpeg", "compression_params": { "quality": 60 } }
},
"monochrome": {
"skip": false,
"downsample": { "skip": false, "downsample_ppi": 300, "threshold_ppi": 450 },
"compression": { "skip": false, "compression_format": "ccitt_g4", "compression_params": {} }
}
},
"fonts": { "subset": false, "compress": false },
"save": { "garbage": 4 }
}
```
## Behavior notes
* Compression results depend on the source PDF. Presets do not promise a fixed reduction, and two levels can produce the same file size.
* When the first compression pass does not produce a smaller PDF, PDF.co retries with the standard configuration while preserving explicit `config` overrides. If the retry is also not smaller, it returns the original PDF.
* Each re-encoded image is kept only when its new stream is smaller. If effective PPI cannot be determined, downsampling is skipped but re-encoding can still run at the original dimensions.
## Responses
A synchronous success returns the final temporary output URL. An asynchronous request returns a `jobId` and a reserved URL that should be used only after the job succeeds — poll it via [Background Job Check](/api/job-check). Errors return `error: true` with a status code and message — see the response examples and [Response Codes](/api/response-codes).
Number of pages in the output PDF.
`false` for a successful request.
PDF.co status code. Success returns `200`. For more information, see [Response Codes](/api/response-codes).
Credits consumed by the request.
Credits remaining for the account.
Processing duration in milliseconds.
Temporary URL of the result. With size protection, it can point to a copy of the original PDF.
Output file name.
UTC timestamp when the temporary URL expires.
Present in the initial response when `async` is `true`.
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
**Sample request**
```bash cURL (minimal) theme={null}
curl --request POST \
--url 'https://api.pdf.co/v2/pdf/compress' \
--header 'Content-Type: application/json' \
--header 'x-api-key: YOUR_API_KEY' \
--data '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf"
}'
```
```bash cURL (preset) theme={null}
curl --request POST \
--url 'https://api.pdf.co/v2/pdf/compress' \
--header 'Content-Type: application/json' \
--header 'x-api-key: YOUR_API_KEY' \
--data '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
"compression_level": "high",
"color_quality": 80
}'
```
```bash cURL (config override) theme={null}
curl --request POST \
--url 'https://api.pdf.co/v2/pdf/compress' \
--header 'Content-Type: application/json' \
--header 'x-api-key: YOUR_API_KEY' \
--data '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
"compression_level": "high",
"config": {
"images": {
"monochrome": {
"compression": {
"compression_format": "ccitt_g4"
}
}
},
"fonts": { "subset": false }
}
}'
```
```javascript Node.js theme={null}
const response = await fetch(
"https://api.pdf.co/v2/pdf/compress",
{
method: "POST",
headers: {
"Content-Type": "application/json",
"x-api-key": "YOUR_API_KEY",
},
body: JSON.stringify({
url: "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
compression_level: "medium",
color_quality: 80,
name: "compressed.pdf",
}),
}
);
const result = await response.json();
if (!response.ok || result.error) {
throw new Error(result.message ?? "PDF.co request failed");
}
console.log(result.url);
```
```python Python theme={null}
import requests
response = requests.post(
"https://api.pdf.co/v2/pdf/compress",
headers={"x-api-key": "YOUR_API_KEY"},
json={
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
"compression_level": "medium",
"color_quality": 80,
"name": "compressed.pdf",
},
)
result = response.json()
if result.get("error"):
raise RuntimeError(result.get("message", "PDF.co request failed"))
print(result["url"])
```
```csharp C# theme={null}
using System.Collections.Generic;
using System.Net.Http;
using System.Text;
using System.Text.Json;
var client = new HttpClient();
client.DefaultRequestHeaders.Add("x-api-key", "YOUR_API_KEY");
var payload = JsonSerializer.Serialize(new Dictionary
{
["url"] = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
["compression_level"] = "medium",
["color_quality"] = 80,
["name"] = "compressed.pdf",
});
var response = await client.PostAsync(
"https://api.pdf.co/v2/pdf/compress",
new StringContent(payload, Encoding.UTF8, "application/json")
);
var json = await response.Content.ReadAsStringAsync();
Console.WriteLine(json);
```
```java Java theme={null}
OkHttpClient client = new OkHttpClient();
String jsonPayload = "{\"url\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf\", \"compression_level\": \"medium\", \"color_quality\": 80, \"name\": \"compressed.pdf\"}";
Request request = new Request.Builder()
.url("https://api.pdf.co/v2/pdf/compress")
.addHeader("x-api-key", "YOUR_API_KEY")
.addHeader("Content-Type", "application/json")
.post(RequestBody.create(MediaType.parse("application/json"), jsonPayload))
.build();
try (Response response = client.newCall(request).execute()) {
System.out.println(response.body().string());
}
```
```php PHP theme={null}
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-compress/sample.pdf",
"compression_level" => "medium",
"color_quality" => 80,
"name" => "compressed.pdf",
]);
$curl = curl_init("https://api.pdf.co/v2/pdf/compress");
curl_setopt_array($curl, [
CURLOPT_HTTPHEADER => [
"x-api-key: YOUR_API_KEY",
"Content-Type: application/json",
],
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_POSTFIELDS => $payload,
]);
echo curl_exec($curl);
curl_close($curl);
```
```json 200 theme={null}
{
"pageCount": 2,
"error": false,
"status": 200,
"credits": 70,
"remainingCredits": 999860,
"duration": 8768,
"url": "https://pdf-temp-files.s3.amazonaws.com/example/sample.pdf",
"name": "sample.pdf",
"outputLinkValidTill": "2026-08-08T12:00:00+00:00"
}
```
```json 400 theme={null}
{
"error": true,
"status": 400,
"message": "Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password)."
}
```
```json 401 theme={null}
{
"error": true,
"status": 401,
"message": "Unauthorized. Authentication is required and has failed or has not yet been provided."
}
```
```json 402 theme={null}
{
"error": true,
"status": 402,
"message": "Not enough credits."
}
```
```json 403 theme={null}
{
"error": true,
"status": 403,
"message": "Access forbidden for input URL."
}
```
```json 404 theme={null}
{
"error": true,
"status": 404,
"message": "The requested resource could not be found."
}
```
```json 408 theme={null}
{
"error": true,
"status": 408,
"message": "The server timed out waiting for the request."
}
```
```json 429 theme={null}
{
"error": true,
"status": 429,
"message": "Too many requests in a given time period."
}
```
```json 441 theme={null}
{
"error": true,
"status": 441,
"message": "Invalid Password. Password protected document."
}
```
```json 442 theme={null}
{
"error": true,
"status": 442,
"message": "Input document is damaged or of incorrect type."
}
```
```json 443 theme={null}
{
"error": true,
"status": 443,
"message": "Permissions. The operation is prohibited by document security settings."
}
```
```json 444 theme={null}
{
"error": true,
"status": 444,
"message": "Profiles parsing error. Please ensure that the configuration is supported."
}
```
```json 445 theme={null}
{
"error": true,
"status": 445,
"message": "Timeout error. For large documents, use asynchronous mode (async=true) and check status via /job/check."
}
```
```json 446 theme={null}
{
"error": true,
"status": 446,
"message": "Some files required for conversion are missing."
}
```
```json 447 theme={null}
{
"error": true,
"status": 447,
"message": "Invalid template."
}
```
```json 448 theme={null}
{
"error": true,
"status": 448,
"message": "Invalid URL or HTML. Ensure the provided URL is valid and accessible."
}
```
```json 449 theme={null}
{
"error": true,
"status": 449,
"message": "Invalid index range. Page index is out of range."
}
```
```json 450 theme={null}
{
"error": true,
"status": 450,
"message": "Invalid page range specified."
}
```
```json 452 theme={null}
{
"error": true,
"status": 452,
"message": "Invalid URL."
}
```
```json 454 theme={null}
{
"error": true,
"status": 454,
"message": "Invalid parameters."
}
```
```json 500 theme={null}
{
"error": true,
"status": 500,
"message": "Something went wrong. Please try again or contact support."
}
```
# PDF Delete Pages
Source: https://developer.pdf.co/api/pdf-delete-pages
Deletes selected pages inside a PDF file.
**Try it live:** [PDF Delete Pages → API Tester](/api-tester/pdf-delete-pages) — send a real request from your browser.
## `POST /v1/pdf/edit/delete-pages`
The `pages` parameter is 1-based, meaning the first page is `1` and not `0`.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *Yes* | - | Specify pages as comma-separated page numbers and ranges to delete (e.g. "1, 2, 5-10" or "3-" for page 3 to the end). The first-page index is 1. Inverted page numbers (e.g. "!1" for the last page) are not supported by this endpoint. This parameter is required: omitting it, or sending an empty or whitespace-only value, returns HTTP 400. The input must be in string format. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf",
"pages": "1-2",
"name": "result.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/d15e5b2c89c04484ae6ac7244ac43ac2/result.pdf",
"pageCount": 2,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 60100
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/delete-pages' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-split/sample.pdf",
"pages": "1-2",
"name": "result.pdf",
"async": false
}'
```
```javascript theme={null}
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require('request');
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/pdf/edit/delete-pages',
'headers': {
'x-api-key': '{{x-api-key}}'
},
formData: {
'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf',
'name': 'result.pdf',
'pages': '1-2'
}
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
```python theme={null}
import requests
url = "https://api.pdf.co/v1/pdf/edit/delete-pages"
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
payload = {'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf',
'name': 'result.pdf',
'pages': '1-2'}
files = [
]
headers = {
'x-api-key': '{{x-api-key}}'
}
response = requests.request("POST", url, headers=headers, json = payload, files = files)
print(response.text.encode('utf8'))
```
```csharp theme={null}
using System;
using RestSharp;
namespace HelloWorldApplication {
class HelloWorld {
static void Main(string[] args) {
var client = new RestClient("https://api.pdf.co/v1/pdf/edit/delete-pages");
client.Timeout = -1;
var request = new RestRequest(Method.POST);
request.AddHeader("x-api-key", "{{x-api-key}}");
request.AlwaysMultipartFormData = true;
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
request.AddParameter("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf");
request.AddParameter("name", "result.pdf");
request.AddParameter("pages", "1-2");
IRestResponse response = client.Execute(request);
Console.WriteLine(response.Content);
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("text/plain");
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM)
.addFormDataPart("url", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf")
.addFormDataPart("name", "result.pdf")
.addFormDataPart("pages", "1-2")
.build();
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/edit/delete-pages")
.method("POST", body)
.addHeader("x-api-key", "{{x-api-key}}")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
"https://api.pdf.co/v1/pdf/edit/delete-pages",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "POST",
CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf','name' => 'result.pdf','pages' => '1-2'),
CURLOPT_HTTPHEADER => array(
"x-api-key: {{x-api-key}}"
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
```
# Extract Attachment
Source: https://developer.pdf.co/api/pdf-extract-attachments
Extracts attachments from a PDF file.
**Try it live:** [Extract Attachment → API Tester](/api-tester/pdf-extract-attachments) — send a real request from your browser.
## `POST /v1/pdf/attachments/extract`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `pageCount` | integer | Number of pages in the PDF document. |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-attachments/attachments.pdf",
"inline": false,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/DO1TAIHEZR5P9QLI7ICYM9DI0AAH57HY/sample.png",
"https://pdf-temp-files.s3.amazonaws.com/EOINIMD7X48JSOB1G8ETLVPOFZLM1NJ2/SampleMetafile.emf",
"https://pdf-temp-files.s3.amazonaws.com/3LW4BXNSPAE0WQTG5DPMXX498OCPNU4Q/ab.tif"
],
"pageCount": 3,
"error": false,
"status": 200,
"name": "attachments.json",
"credits": 24,
"duration": 1211,
"remainingCredits": 98003902
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/attachments/extract' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-attachments/attachments.pdf",
"inline": false,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var fs = require("fs");
var path = require("path");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf";
// Prepare request for API endpoint
var queryPath = `/v1/pdf/attachments/extract`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
let responseData = '';
response.setEncoding("utf8");
response.on("data", (chunk) => {
responseData += chunk;
});
response.on("end", () => {
// Parse JSON response
var data = JSON.parse(responseData);
if (data.error == false) {
// Download extracted files
data.urls.forEach((url) => {
var localFileName = path.basename(url);
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated file saved as "${localFileName}" file.`);
});
});
}, this);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.error(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import requests
import json
url = "https://api.pdf.co/v1/pdf/attachments/extract"
payload = json.dumps({
"url": "https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf",
"inline": True,
"async": False
})
headers = {
'Content-Type': 'application/json',
'x-api-key': '__Replace_With_Your_PDFco_API_Key__'
}
response = requests.request("POST", url, headers=headers, data=payload)
print(response.text)
```
```csharp theme={null}
using System;
using RestSharp;
namespace HelloWorldApplication {
class HelloWorld {
static void Main(string[] args) {
var client = new RestClient("https://api.pdf.co/v1/pdf/attachments/extract");
client.Timeout = -1;
var request = new RestRequest(Method.POST);
request.AddHeader("Content-Type", "application/json");
request.AddHeader("x-api-key", "__Replace_With_Your_PDFco_API_Key__");
var body = @"{" + "\n" +
@" ""url"": ""https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf""," + "\n" +
@" ""inline"": true," + "\n" +
@" ""async"": false" + "\n" +
@"}";
request.AddParameter("application/json", body, ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
Console.WriteLine(response.Content);
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("application/json");
RequestBody body = RequestBody.create(mediaType, "{\n \"url\": \"https://bytescout-com.s3.us-west-2.amazonaws.com/files/demo-files/cloud-api/pdf-attachments/attachments.pdf\",\n \"inline\": true,\n \"async\": false\n}");
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/attachments/extract")
.method("POST", body)
.addHeader("Content-Type", "application/json")
.addHeader("x-api-key", "__Replace_With_Your_PDFco_API_Key__")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
Cloud API asynchronous "Extract PDF Attachment" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF Find Text
Source: https://developer.pdf.co/api/pdf-find/basic
Find text in PDF and get coordinates. Supports regular expressions.
**Try it live:** [PDF Find Text → API Tester](/api-tester/pdf-find/basic) — send a real request from your browser.
## `POST /v1/pdf/find`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`.
| Attribute | Type | Required | Default | Description |
| ------------------------------------------------------ | ------- | -------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `searchString` | string | *Yes* | - | Text to search can support regular expressions if you set the `regexSearch` param to true. |
| `wordMatchingMode` | string | *No* | None | WordMatchingMode defines how search terms match PDF text. Modes: `None` (exact string match only), `SmartMatch` (default; flexible word boundary match, includes letters/digits/punctuation), `ExactMatch` (strict word boundaries, whole-word match only). |
| `regexSearch` | boolean | *No* | `false` | Set to true to enable regular expression search for the `searchString(s)` parameter. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `ColumnDetectionMode` | string | *No* | ContentGroupsAndBorders | Controls column detection/alignment in PDF table extraction. Modes: `ContentGroupsAndBorders` (default; text + lines), `ContentGroups` (text grouping only), `Borders` (lines only), `BorderedTables` (OCR-based for bordered tables), `ContentGroupsAI` (AI for dense/complex layouts). |
| `DetectionMinNumberOfRows` | integer | *No* | 1 | Minimum number of rows to detect in a table |
| `DetectionMinNumberOfColumns` | integer | *No* | 1 | Minimum number of columns to detect in a table |
| `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | *No* | `0` | Maximum number of invalid subsequent rows allowed in a table |
| `DetectionMinNumberOfLineBreaksBetweenTables` | integer | *No* | `0` | Minimum number of line breaks between tables |
| `EnhanceTableBorders` | boolean | *No* | `true` | Enhance table borders or not |
| `OCRDetectPageRotation` | boolean | *No* | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See [Support page rotation](#support-page-rotation) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `requestParametersDocument` | string | *No* | - | |
| `responseParameters` | object | *No* | - | - |
| `error` | boolean | *No* | - | Indicates whether an error occurred (`false` means success) |
| `status` | string | *No* | - | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `message` | string | *No* | - | Message of the request |
| `credits` | integer | *No* | - | Number of credits consumed by the request |
| `remainingCredits` | integer | *No* | - | Number of credits remaining in the account |
| `duration` | integer | *No* | - | Time taken for the operation in milliseconds |
| `errorCode` | integer | *No* | - | Error code of the request (400, 401, 402, 403, 404, 500, etc.) |
### Support page rotation
This endpoint supports **PDF** page rotation as follows:
```json theme={null}
{
"profiles": "{ 'OCRDetectPageRotation': true }"
}
```
### Find only bordered tables
You can limit search to bordered tables only by enabling the *legacy table* search mode with the following `profiles` config:
```json theme={null}
{
"profiles": "{ 'Mode': 'Legacy',
'ColumnDetectionMode': 'BorderedTables',
'DetectionMinNumberOfRows': 1,
'DetectionMinNumberOfColumns': 1,
'DetectionMaxNumberOfInvalidSubsequentRowsAllowed': 0,
'DetectionMinNumberOfLineBreaksBetweenTables': 0,
'EnhanceTableBorders': false
}"
}
```
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"async": "false",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"searchString": "Invoice Date \\d+/\\d+/\\d+",
"regexSearch": "true",
"name": "output",
"pages": "0-",
"inline": "true",
"wordMatchingMode": "",
"password": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": [
{
"text": "Invoice Date 01/01/2016",
"left": 436.5400085449219,
"top": 130.4599995137751,
"width": 122.85311957550027,
"height": 11.040000486224898,
"pageIndex": 0,
"bounds": {
"location": {
"isEmpty": false,
"x": 436.54,
"y": 130.46
},
"size": "122.853119, 11.0400009",
"x": 436.54,
"y": 130.46,
"width": 122.853119,
"height": 11.0400009,
"left": 436.54,
"top": 130.46,
"right": 559.3931,
"bottom": 141.5,
"isEmpty": false
},
"elementCount": 1,
"elements": [
{
"index": 0,
"left": 436.5400085449219,
"top": 130.4599995137751,
"width": 122.85311957550027,
"height": 11.040000486224898,
"angle": 0,
"text": "Invoice Date 01/01/2016",
"isNewLine": true,
"fontIsBold": true,
"fontIsItalic": false,
"fontName": "Helvetica-Bold",
"fontSize": 11,
"fontColor": "0, 0, 0",
"fontColorAsOleColor": 0,
"fontColorAsHtmlColor": "#000000",
"bounds": {
"location": {
"isEmpty": false,
"x": 436.54,
"y": 130.46
},
"size": "122.853119, 11.0400009",
"x": 436.54,
"y": 130.46,
"width": 122.853119,
"height": 11.0400009,
"left": 436.54,
"top": 130.46,
"right": 559.3931,
"bottom": 141.5,
"isEmpty": false
}
}
]
}
],
"pageCount": 1,
"error": false,
"status": 200,
"name": "output",
"remainingCredits": 59970
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/find' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"async": "false",
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"searchString": "Invoice Date \\d+/\\d+/\\d+",
"regexSearch": "true",
"name": "output",
"pages": "0-",
"inline": "true",
"wordMatchingMode": "",
"password": ""
}'
```
```javascript theme={null}
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Search string.
const SearchString = '[4-9][0-9].[0-9][0-9]'; // Regular expression to find numbers in format dd.dd and between 40.00 to 99.99
// Enable regular expressions (Regex)
const RegexSearch = 'True';
// Prepare URL for PDF text search API call.
// See documentation: https://developer.pdf.co
var query = `https://api.pdf.co/v1/pdf/find`;
let reqOptions = {
uri: query,
headers: { "x-api-key": API_KEY },
formData: {
password: Password,
pages: Pages,
url: SourceFileUrl,
searchString: SearchString,
regexSearch: RegexSearch
}
};
// Send request
request.post(reqOptions, function (error, response, body) {
if (error) {
return console.error("Error: ", error);
}
// Parse JSON response
let data = JSON.parse(body);
for (let index = 0; index < data.body.length; index++) {
const element = data.body[index];
console.log("Found text " + element["text"] + " at coordinates " + element["left"] + ", " + element["top"]);
}
});
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Search string.
SearchString = "\d{1,}\.\d\d" # Regular expression to find numbers like '100.00'
# Note: do not use `+` char in regex, but use `{1,}` instead.
# `+` char is valid for URL and will not be escaped, and it will become a space char on the server side.
# Enable regular expressions (Regex)
RegexSearch = True
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
searchTextInPDF(uploadedFileUrl)
def searchTextInPDF(uploadedFileUrl):
"""Search Text using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {}
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
parameters["searchString"] = SearchString
parameters["regexSearch"] = RegexSearch
# Prepare URL for 'PDF Text Search' API request
url = "{}/pdf/find".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Display found information
for item in json["body"]:
print(f"Found text {item['text']} at coordinates {item['left']}, {item['top']}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Search string.
const string SearchString = @"\d{1,}\.\d\d"; // Regular expression to find numbers like '100.00'
// Note: do not use `+` char in regex, but use `{1,}` instead.
// `+` char is valid for URL and will not be escaped, and it will become a space char on the server side.
// Enable regular expressions (Regex)
const bool RegexSearch = true;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. MAKE UPLOADED PDF FILE SEARCHABLE
// URL for `PDF Text Search` API call
// See documentation: https://developer.pdf.co
string url = "https://api.pdf.co/v1/pdf/find";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
parameters.Add("searchString", SearchString);
parameters.Add("regexSearch", RegexSearch);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
foreach (JToken item in json["body"])
{
Console.WriteLine($"Found text \"{item["text"]}\" at coordinates {item["left"]}, {item["top"]}");
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException ex)
{
Console.WriteLine(ex.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Search string.
final static String SearchString = "\\d{1,}\\.\\d\\d"; // Regular expression to find numbers like '100.00'
// Note: do not use `+` char in regex, but use `{1,}` instead.
// `+` char is valid for URL and will not be escaped, and it will become a space char on the server side.
// Enable regular expressions (Regex)
final static boolean RegexSearch = true;
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for PDF text search API call.
// See documentation: https://developer.pdf.co
String query = "https://api.pdf.co/v1/pdf/find";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\", \"searchString\": \"%s\", \"regexSearch\": \"%s\"}",
Password,
Pages,
SourceFileURL,
SearchString,
RegexSearch);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Display found items in console
for (JsonElement element : json.get("body").getAsJsonArray())
{
JsonObject item = (JsonObject) element;
System.out.println("Found text " + item.get("text") + " at coordinates " + item.get("left") + ", "+ item.get("top"));
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
}
```
# Find Text in Table with AI
Source: https://developer.pdf.co/api/pdf-find/table
Detect tables in PDF and scanned documents using AI, returning the page, coordinates, and detected column structure for each table.
**Try it live:** [Find Text in Table with AI → API Tester](/api-tester/pdf-find/table) — send a real request from your browser.
## `POST /v1/pdf/find/table`
This function finds tables in documents using an AI-powered table detection engine.
This endpoint locates `tables` in an input PDF document and returns JSON with:
* The array of tables objects.
* `X`, `Y`, `Width`, and `Height` coordinates for every table found.
* `Rect` param for every table that you can re-use with `pdf/convert/to/json`, `pdf/convert/to/csv`, `pdf/convert/to/csv`, and other endpoints to extract a selected table only.
* `PageIndex` page index for a page with a table. The very first page is `0` .
* `Columns` array with the set of `X` coordinates for every column inside the table that was found.
To extract the table into CSV, JSON, or XML please use pdf/convert/to/csv, pdf/convert/to/json2, and pdf/convert/to/xml endpoints with rect parameter value from rect output param for this table accordingly.
To extract the table into CSV, JSON, or XML please use [pdf/convert/to/csv](/api/pdf-to-csv), [pdf/convert/to/json2](/api/pdf-to-json/with-ai), and [pdf/convert/to/xml](/api/pdf-to-xml) endpoints with `rect` parameter value from `rect` output param for this table accordingly.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Find only bordered tables
You can limit search to bordered tables only by enabling the *legacy table* search mode with the following `profiles` config:
```json theme={null}
{
"profiles": "{ 'Mode': 'Legacy',
'ColumnDetectionMode': 'BorderedTables',
'DetectionMinNumberOfRows': 1,
'DetectionMinNumberOfColumns': 1,
'DetectionMaxNumberOfInvalidSubsequentRowsAllowed': 0,
'DetectionMinNumberOfLineBreaksBetweenTables': 0,
'EnhanceTableBorders': false
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `body` | object | Response body. |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"async": "false",
"inline": "true",
"password": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"tables": [
{
"PageIndex": 0,
"X": 36,
"Y": 34.4400024,
"Width": 523.44,
"Height": 160.82,
"Columns": [
357.675
],
"rect": "36, 34.4400024, 523.44, 160.82"
},
{
"PageIndex": 0,
"X": 36,
"Y": 316.249969,
"Width": 523.44,
"Height": 120.620026,
"Columns": [
157.117,
340.68,
475.84
],
"rect": "36, 316.249969, 523.44, 120.620026"
}
]
},
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.json",
"remainingCredits": 98892697,
"credits": 21
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/find/table' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "pdfco-test-files.s3.us-west-2.amazonaws.compdf-to-text/sample.pdf",
"async": "false",
"inline": "true",
"password": ""
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Prepare URL for PDF Table Search API call.
// See documentation: https://developer.pdf.co
var query = `https://api.pdf.co/v1/pdf/find/table`;
let reqOptions = {
uri: query,
headers: { "x-api-key": API_KEY },
formData: {
password: Password,
pages: Pages,
url: SourceFileUrl
}
};
// Send request
request.post(reqOptions, function (error, resp, body) {
if (error) {
return console.error("Error: ", error);
}
var jsonBody = JSON.parse(body);
// Loop through all found tables, and get json data
if (jsonBody.body.tables && jsonBody.body.tables.length > 0) {
for (var i = 0; i < jsonBody.body.tables.length; i++) {
getJSONFromCoordinates(SourceFileUrl, jsonBody.body.tables[i].PageIndex, jsonBody.body.tables[i].rect, `table_${i + 1}.json`);
}
}
});
/**
* Get JSON from specific co-ordinates
*/
function getJSONFromCoordinates(fileUrl, pageIndex, rect, outputFileName) {
// Prepare request to `PDF To JSON` API endpoint
var jsonQueryPath = `https://api.pdf.co/v1/pdf/convert/to/json`;
// Json Request
let jsonReqOptions = {
uri: jsonQueryPath,
headers: { "x-api-key": API_KEY },
formData: {
pages: pageIndex,
url: fileUrl,
rect: rect
}
};
// Send request
request.post(jsonReqOptions, function (error, resp, body) {
if (error) {
return console.error("Error: ", error);
}
var outputJsonUrl = JSON.parse(body).url;
// Download JSON file
var file = fs.createWriteStream(outputFileName);
https.get(outputJsonUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated JSON file saved as "${outputFileName}" file.`);
});
});
});
}
```
```python theme={null}
import requests
import os
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***************************************"
# Direct URL of source PDF file.
SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Prepare URL for PDF Table Search API call.
query = "https://api.pdf.co/v1/pdf/find/table"
reqOptions = {
'password': Password,
'pages': Pages,
'url': SourceFileUrl
}
headers = {
'x-api-key': API_KEY
}
def getJSONFromCoordinates(fileUrl, pageIndex, rect, outputFileName):
# Prepare request to `PDF To JSON` API endpoint
jsonQueryPath = "https://api.pdf.co/v1/pdf/convert/to/json"
# Json Request
jsonReqOptions = {
'pages': pageIndex,
'url': fileUrl,
'rect': rect
}
# Send request
response = requests.post(jsonQueryPath, headers=headers, data=jsonReqOptions)
if response.status_code == 200:
outputJsonUrl = response.json()['url']
# Download JSON file
res = requests.get(outputJsonUrl)
with open(outputFileName, 'wb') as outfile:
outfile.write(res.content)
print(f'Generated JSON file saved as "{outputFileName}" file.')
else:
print(f"Request error: {response.status_code} {response.reason}")
# Send request
response = requests.post(query, headers=headers, data=reqOptions)
if response.status_code == 200:
jsonBody = response.json()
# Loop through all found tables, and get json data
if 'tables' in jsonBody['body'] and len(jsonBody['body']['tables']) > 0:
for i, table in enumerate(jsonBody['body']['tables']):
getJSONFromCoordinates(SourceFileUrl, table['PageIndex'], table['rect'], f"table_{i + 1}.json")
else:
print(f"Request error: {response.status_code} {response.reason}")
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.Net;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*****************************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for PDF Table Search API call.
// See documentation: https://developer.pdf.co
string url = "https://api.pdf.co/v1/pdf/find/table";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", SourceFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["status"].ToString() != "error")
{
Console.WriteLine(response);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for PDF Table Search API call.
// See documentation: https://developer.pdf.co
String query = "https://api.pdf.co/v1/pdf/find/table";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
Password,
Pages,
SourceFileURL);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
System.out.println(response.body().string());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
}
```
# PDF from CSV
Source: https://developer.pdf.co/api/pdf-from-document/csv
Convert `CSV`, `XLS`, `XLSX` files into `PDF`.
**Try it live:** [PDF from CSV → API Tester](/api-tester/pdf-from-document/csv) — send a real request from your browser.
## `POST /pdf/convert/from/csv`
During conversion you should not expect any Word macros to operate as we do not support Office macros.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `autosize` | boolean | *No* | - | Set to `true` to page dimensions adjust to content with automatic page sizing. If false, uses worksheet's page setup. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"pages": "0-",
"name": "result.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/efc283805b4a47da87910826d4ddf063/result.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEKz%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDFzXkfTapcUbLLKahiKBAbIL4F2wV3gvozuGDxmOpWUu9ETuVzkYKjMuNLAFzVZeSgRm9Yuaj7ubad9uOLQkL65GNgBQoy1Xm%2FxtLWD9tegUYd3hFvYfIWMfkWjuROwMGTZeD3CMacDPdFkP%2BUSG4aXOZb8MoG2PXnsd9UUeOvrevZkCVTg77OBXIteBCPOojSjeis%2F5BTIoVCi%2FrwV5kEGkbfBwtgsfQL3MxSbg7j%2Fud%2F3oGbUWW7zsemcfiHTiFg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAK44PZ6O/20200812/us-west-2/s3/aws4_request&X-Amz-Date=20200812T103301Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=6a176281828de74c917a4ff5bacd46eeca50221178bf34d01cac331f172e51c3",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 61165
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/doc' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"pages": "0-",
"name": "result.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `DOC to PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/doc`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source DOC file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
convertDOCToPDF(SourceFileURL, DestinationFile)
def convertDOCToPDF(uploadedFileUrl, destinationFile):
"""Converts DOC to PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrl
# Prepare URL for 'DOC To PDF' API request
url = "{}/pdf/convert/from/doc".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", SourceFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// URL of `DOC To PDF` API call
string url = "https://api.pdf.co/v1/pdf/convert/from/doc";
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `DOC To PDF` API call
String query = "https://api.pdf.co/v1/pdf/convert/from/doc";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "DOC To PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF from DOC
Source: https://developer.pdf.co/api/pdf-from-document/doc
Convert `DOC`, `DOCX`, `RTF`, `TXT`, `XPS`, `PPT`, `PPTX` files into `PDF`.
**Try it live:** [PDF from DOC → API Tester](/api-tester/pdf-from-document/doc) — send a real request from your browser.
## `POST /pdf/convert/from/doc`
During conversion you should not expect any Word macros to operate as we do not support Office macros.This endpoint can be utilized as-is to convert PowerPoint files (`PPT`, `PPTX`) to PDF. However, `PPT` and `PPTX` **are not natively supported**, and we may not provide assistance or fixes for any issues, bugs, or limitations encountered during its use.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `autosize` | boolean | *No* | - | Set to `true` to page dimensions adjust to content with automatic page sizing. If false, uses worksheet's page setup. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"pages": "0-",
"name": "result.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3-us-west-2.amazonaws.com/efc283805b4a47da87910826d4ddf063/result.pdf?X-Amz-Expires=3600&x-amz-security-token=FwoGZXIvYXdzEKz%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDFzXkfTapcUbLLKahiKBAbIL4F2wV3gvozuGDxmOpWUu9ETuVzkYKjMuNLAFzVZeSgRm9Yuaj7ubad9uOLQkL65GNgBQoy1Xm%2FxtLWD9tegUYd3hFvYfIWMfkWjuROwMGTZeD3CMacDPdFkP%2BUSG4aXOZb8MoG2PXnsd9UUeOvrevZkCVTg77OBXIteBCPOojSjeis%2F5BTIoVCi%2FrwV5kEGkbfBwtgsfQL3MxSbg7j%2Fud%2F3oGbUWW7zsemcfiHTiFg%3D%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHAK44PZ6O/20200812/us-west-2/s3/aws4_request&X-Amz-Date=20200812T103301Z&X-Amz-SignedHeaders=host;x-amz-security-token&X-Amz-Signature=6a176281828de74c917a4ff5bacd46eeca50221178bf34d01cac331f172e51c3",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 61165
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/doc' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx",
"pages": "0-",
"name": "result.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `DOC to PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/doc`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source DOC file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
convertDOCToPDF(SourceFileURL, DestinationFile)
def convertDOCToPDF(uploadedFileUrl, destinationFile):
"""Converts DOC to PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrl
# Prepare URL for 'DOC To PDF' API request
url = "{}/pdf/convert/from/doc".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", SourceFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// URL of `DOC To PDF` API call
string url = "https://api.pdf.co/v1/pdf/convert/from/doc";
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source DOC or DOCX file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/doc-to-pdf/sample.docx";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `DOC To PDF` API call
String query = "https://api.pdf.co/v1/pdf/convert/from/doc";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "DOC To PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF from Email
Source: https://developer.pdf.co/api/pdf-from-email
Convert email files (`.msg` or `.eml`) code into PDF. Extract attachments (if any) from input email and embeds into PDF as PDF attachments.
**Try it live:** [PDF from Email → API Tester](/api-tester/pdf-from-email) — send a real request from your browser.
## `POST /v1/pdf/convert/from/email`
Images and attachments within `.eml` and `.msg` files **must be publicly accessible.** Resources stored on local file systems or gated behind authentication are not supported for processing.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| --------------------------------------------------------- | ------- | -------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. |
| `paperSize` | string | *No* | `A4` | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. |
| `orientation` | string | *No* | `Portrait` | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `embedAttachments` | boolean | *No* | `true` | Set to true to automatically embeds all attachments from original input email MSG or EML files into the final output PDF. Set it to false if you don't want to embed attachments so it will convert only the body of the input email. True by default. |
| `convertAttachments` | boolean | *No* | `true` | Set to false if you don't want to convert attachments from the original email and want to embed them as original files (as embedded PDF attachments). Converts attachments that are supported into PDF format and then merges into output final PDF. The supported attachment types for conversion are: `.eml`, `.html/.htm`, `.pdf`, `.doc`, `.docx`, and `.rtf`. Non-supported file types are added as PDF attachments (Adobe Reader or another viewer may be required to view PDF attachments). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| [`removeHTMLHeadStyleTags`](#removehtmlheadstyletags) | string | *No* | `false` | Removes default styles from the `` section of Outlook emails to ensure accurate PDF rendering. Set to `'true'` to enable. |
| [`removeHTMLBodyStyleTags`](#removehtmlbodystyletags) | string | *No* | `false` | Removes inline and default styles from the `` section of Outlook emails to ensure accurate PDF rendering. Set to `'true'` to enable. |
| [`CustomScript`](#customscript) | string | *No* | - | Custom JavaScript code executed on the email content before PDF conversion. Use to modify HTML elements, disable links, or apply custom transformations. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Profiles examples
#### removeHTMLHeadStyleTags
* **Remove default Outlook styles for accurate PDF conversion**
```json theme={null}
{
"profiles": {
"removeHTMLHeadStyleTags": "true",
}
}
```
#### removeHTMLBodyStyleTags
* **Remove inline and default styles for accurate PDF conversion**
```json theme={null}
{
"profiles": {
"removeHTMLBodyStyleTags": "true"
}
}
```
This example removes all inline and default styles from the `` section of Outlook emails to ensure accurate PDF rendering.
#### CustomScript
```json theme={null}
{
"profiles": {
"CustomScript": "document.querySelectorAll('a').forEach(a => { a.href = '#' });"
}
}
```
This example disables all active links while keeping the link text visible.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml",
"embedAttachments": true,
"convertAttachments": true,
"paperSize": "Letter",
"name": "email-with-attachments",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/980bc13f061344809c75e83ce181851c/Contact_us.pdf",
"pageCount": 3,
"error": false,
"status": 200,
"name": "Contact_us.pdf",
"remainingCredits": 60637
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/email' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml",
"embedAttachments": true,
"convertAttachments": true,
"paperSize": "Letter",
"name": "email-with-attachments",
"async": false
}'
```
```javascript theme={null}
var request = require('request');
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
var options = {
'method': 'POST',
'url': 'https://api.pdf.co/v1/pdf/convert/from/email',
'headers': {
'Content-Type': 'application/json',
'x-api-key': '{{x-api-key}}'
},
formData: {
'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml',
'embedAttachments': 'true',
'convertAttachments': 'true',
'paperSize': 'Letter',
'name': 'email-with-attachments',
'async': 'false'
}
};
request(options, function (error, response) {
if (error) throw new Error(error);
console.log(response.body);
});
```
```python theme={null}
import requests
url = "https://api.pdf.co/v1/pdf/convert/from/email"
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
payload={'url': 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml',
'embedAttachments': 'true',
'convertAttachments': 'true',
'paperSize': 'Letter',
'name': 'email-with-attachments',
'async': 'false'}
files=[
]
headers = {
'Content-Type': 'application/json',
'x-api-key': '{{x-api-key}}'
}
response = requests.request("POST", url, headers=headers, json=payload, files=files)
print(response.text)
```
```csharp theme={null}
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
using System;
using System.Collections.Generic;
using System.Net;
using System.Threading;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source Email file to convert
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = @"https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml";
// Ouput file path
const string DestinationFile = @"output.pdf";
// (!) Make asynchronous job
const bool Async = true;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
try
{
// URL for `PDF FROM Email` API call
var url = "https://api.pdf.co/v1/pdf/convert/from/email";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
// Link to input EML or MSG file to be converted.
// You can pass link to file from Google Drive, Dropbox or another online file service that can generate shareable links.
// You can also use built-in PDF.co cloud storage located at https://app.pdf.co/files or upload your file as temporary file right before making this API call (see Upload and Manage Files section for more details on uploading files via API).
parameters.Add("url", SourceFileUrl);
// True by default.
// Set to true to automatically embeds all attachments from original input email MSG or EML fileas files into final output PDF.
// Set to false if you don’t want to embed attachments so it will convert only the body of input email.
parameters.Add("embedAttachments", true);
// true by default.
// Converts attachments that are supported by API (doc, docx, html, png, jpg etc) into PDF and merges into output final PDF.
// Non-supported file types are added as PDF attachments (Adobe Reader or another viewer maybe required to view PDF attachments).
// Set to false if you don’t want to convert attachments from original email and want to embed them as original files (as embedded pdf attachments).
parameters.Add("convertAttachments", true);
// Can be Letter, A4, A5, A6 or custom size like 200x200
parameters.Add("paperSize", "Letter");
// Name of output PDF
parameters.Add("name", "email-with-attachments");
// Set to true to run as async job in background (recommended for heavy documents).
parameters.Add("async", Async);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Asynchronous job ID
string jobId = json["jobId"].ToString();
// URL of generated JSON file available after the job completion; it will contain URLs of result PDF files.
string resultFileUrl = json["url"].ToString();
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do
{
string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status);
if (status == "success")
{
// Download output file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
break;
}
else if (status == "working")
{
// Pause for a few seconds
Thread.Sleep(3000);
}
else
{
Console.WriteLine(status);
break;
}
}
while (true);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
///
/// Checks Job Status
///
static string CheckJobStatus(string jobId)
{
using (WebClient webClient = new WebClient())
{
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
string response = webClient.DownloadString(url);
JObject json = JObject.Parse(response);
return Convert.ToString(json["status"]);
}
}
}
}
```
```java theme={null}
import java.io.*;
import okhttp3.*;
public class main {
public static void main(String []args) throws IOException{
OkHttpClient client = new OkHttpClient().newBuilder()
.build();
MediaType mediaType = MediaType.parse("application/json");
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
RequestBody body = new MultipartBody.Builder().setType(MultipartBody.FORM)
.addFormDataPart("url","https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml")
.addFormDataPart("embedAttachments","true")
.addFormDataPart("convertAttachments","true")
.addFormDataPart("paperSize","Letter")
.addFormDataPart("name","email-with-attachments")
.addFormDataPart("async","false")
.build();
Request request = new Request.Builder()
.url("https://api.pdf.co/v1/pdf/convert/from/email")
.method("POST", body)
.addHeader("Content-Type", "application/json")
.addHeader("x-api-key", "{{x-api-key}}")
.build();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
}
}
```
```php theme={null}
'https://api.pdf.co/v1/pdf/convert/from/email',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => '',
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 0,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_POSTFIELDS => array('url' => 'https://pdfco-test-files.s3.us-west-2.amazonaws.com/email-to-pdf/sample.eml','embedAttachments' => 'true','convertAttachments' => 'true','paperSize' => 'Letter','name' => 'email-with-attachments','async' => 'false'),
CURLOPT_HTTPHEADER => array(
'Content-Type: application/json',
'x-api-key: {{x-api-key}}'
),
));
$response = json_decode(curl_exec($curl));
curl_close($curl);
echo "
Output:
", var_export($response, true), "
";
```
# PDF from HTML
Source: https://developer.pdf.co/api/pdf-from-html/convert
Convert raw HTML markup into a PDF document with custom margins, paper size, orientation, headers and footers.
**Try it live:** [PDF from HTML → API Tester](/api-tester/pdf-from-html/convert) — send a real request from your browser.
## `POST /v1/pdf/convert/from/html`
This API converts a RAW HTML code into a PDF document and process any JavaScript which the webpage triggers when it loads. For example if the the webpage triggers a JavaScript popup window then that will be included in the conversion process. There is no option to disable JavaScript on the supplied RAW HTML code.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`Remember to ensure that request sizes are less than `4` mb in file size. For more information, see [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `html` | string | *Yes* | - | Input HTML code to be converted. To convert the link to a PDF use the /pdf/convert/from/url endpoint instead. |
| `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. |
| `paperSize` | string | *No* | A4 | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. |
| `orientation` | string | *No* | Portrait | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. |
| `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL |
| `mediaType` | string | *No* | print | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). |
| `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. |
| `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. |
| `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Header & Footer
The `header` and `footer` parameters can contain valid HTML markup with the following classes used to inject printing values into them:
* `date`: formatted print date
* `title`: document title
* `url`: document location
* `pageNumber`: current page number
* `totalPages`: total pages in the document
img: tag is supported in both the header and footer parameter, provided that the `src` attribute is specified as a `base64-encoded` string.
For example, the following markup will generate `Page N of NN` page numbering:
```html theme={null}
Page of .
```
### Sample Header & Footer
An example with an advanced `header` and `footer`. Note that the top and bottom page margins are important because page content may overlap the footer or header.
```json theme={null}
{
"html": "
"
}
```
If you use `JSON` as input then make sure to escape it first (with `JSON.stringify(dataObject)` in JS). Escaping is when every `"` is replaced with `\"`. Example with `"` be escaped as `\"` then: `"templateData": "{ \"paid\": true, \"invoice_id\": \"0002\", \"total\": \"$999.99\" }"`.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"html": "
Hello World!
Go to PDF.co",
"name": "result.pdf",
"margins": "5px 5px 5px 5px",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"mediaType": "print",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 60646
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/html' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"html": "
Hello World!
Go to PDF.co",
"name": "result.pdf",
"margins": "5px 5px 5px 5px",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"mediaType": "print",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***************************";
// HTML Input
const inputHtml = "./sample.html";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare requests params as JSON
var parameters = {};
// Input HTML code to be converted. Required.
parameters["html"] = fs.readFileSync(inputHtml, "utf8");
// Name of resulting file
parameters["name"] = path.basename(DestinationFile);
// Set to css style margins like 10 px or 5px 5px 5px 5px.
parameters["margins"] = "5px 5px 5px 5px";
// Can be Letter, A4, A5, A6 or custom size like 200x200
parameters["paperSize"] = "Letter";
// Set to Portrait or Landscape. Portrait by default.
parameters["orientation"] = "Portrait";
// true by default. Set to false to disbale printing of background.
parameters["printBackground"] = true;
// If large input document, process in async mode by passing true
parameters["async"] = false;
// Set to HTML for header to be applied on every page at the header.
parameters["header"] = "";
// Set to HTML for footer to be applied on every page at the bottom.
parameters["footer"] = "";
// Convert JSON object to string
var jsonPayload = JSON.stringify(parameters);
// Prepare request to `HTML To PDF` API endpoint
var url = '/v1/pdf/convert/from/html';
var reqOptions = {
host: "api.pdf.co",
path: url,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import json
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "**************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# HTML template
file_read = open(".\\sample.html", mode='r', encoding= 'utf-8')
SampleHtml = file_read.read()
file_read.close()
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
GeneratePDFFromHtml(SampleHtml, DestinationFile)
def GeneratePDFFromHtml(SampleHtml, destinationFile):
"""Converts HTML to PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
# Input HTML code to be converted. Required.
parameters["html"] = SampleHtml
# Name of resulting file
parameters["name"] = os.path.basename(destinationFile)
# Set to css style margins like 10 px or 5px 5px 5px 5px.
parameters["margins"] = "5px 5px 5px 5px"
# Can be Letter, A4, A5, A6 or custom size like 200x200
parameters["paperSize"] = "Letter"
# Set to Portrait or Landscape. Portrait by default.
parameters["orientation"] = "Portrait"
# true by default. Set to false to disable printing of background.
parameters["printBackground"] = "true"
# If large input document, process in async mode by passing true
parameters["async"] = "false"
# Set to HTML for header to be applied on every page at the header.
parameters["header"] = ""
# Set to HTML for footer to be applied on every page at the bottom.
parameters["footer"] = ""
# Prepare URL for 'HTML To PDF' API request
url = "{}/pdf/convert/from/html".format(
BASE_URL
)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "******************************";
static void Main(string[] args)
{
// HTML input
string inputSample = File.ReadAllText(@".\sample.html");
// Destination PDF file name
string destinationFile = @".\result.pdf";
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Set JSON content type
webClient.Headers.Add("Content-Type", "application/json");
try
{
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
// Input HTML code to be converted. Required.
parameters.Add("html", inputSample);
// Name of resulting file
parameters.Add("name", Path.GetFileName(destinationFile));
// Set to css style margins like 10 px or 5px 5px 5px 5px.
parameters.Add("margins", "5px 5px 5px 5px");
// Can be Letter, A4, A5, A6 or custom size like 200x200
parameters.Add("paperSize", "Letter");
// Set to Portrait or Landscape. Portrait by default.
parameters.Add("orientation", "Portrait");
// true by default. Set to false to disbale printing of background.
parameters.Add("printBackground", true);
// If large input document, process in async mode by passing true
parameters.Add("async", false);
// Set to HTML for header to be applied on every page at the header.
parameters.Add("header", "");
// Set to HTML for footer to be applied on every page at the bottom.
parameters.Add("footer", "");
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Prepare URL for `HTML to PDF` API call
string url = "https://api.pdf.co/v1/pdf/convert/from/html";
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
webClient.Headers.Remove("Content-Type"); // remove the header required for only the previous request
// Download the PDF file
webClient.DownloadFile(resultFileUrl, destinationFile);
Console.WriteLine("Generated PDF document saved as \"{0}\" file.", destinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key to exit...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
// HTML input
final String inputSample = new String(Files.readAllBytes(Paths.get(".\\sample.html")));
// Destination PDF file name
final Path destinationFile = Paths.get(".\\result.pdf");
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `HTML to PDF` API call
String apiUrl = "https://api.pdf.co/v1/pdf/convert/from/html";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, apiUrl, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Prepare request body in JSON format
JsonObject jsonBody = new JsonObject();
// Input HTML code to be converted. Required.
jsonBody.add("html", new JsonPrimitive(inputSample));
// Name of resulting file
jsonBody.add("name", new JsonPrimitive(destinationFile.getFileName().toString()));
// Set to css style margins like 10 px or 5px 5px 5px 5px.
jsonBody.add("margins", new JsonPrimitive("5px 5px 5px 5px"));
// Can be Letter, A4, A5, A6 or custom size like 200x200
jsonBody.add("paperSize", new JsonPrimitive("Letter"));
// Set to Portrait or Landscape. Portrait by default.
jsonBody.add("orientation", new JsonPrimitive("Portrait"));
// true by default. Set to false to disable printing of background.
jsonBody.add("printBackground", new JsonPrimitive(true));
// If large input document, process in async mode by passing true
jsonBody.add("async", new JsonPrimitive(false));
// Set to HTML for header to be applied on every page at the header.
jsonBody.add("header", new JsonPrimitive(""));
// Set to HTML for footer to be applied on every page at the bottom.
jsonBody.add("footer", new JsonPrimitive(""));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
HTML to PDF Result
Error: " . $json["message"] . "";
}
else
{
$resultFileUrl = $json["url"];
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***************************";
// HTML Input
const inputHtml = "./sample.html";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare requests params as JSON
var parameters = {};
// Input HTML code to be converted. Required.
parameters["html"] = fs.readFileSync(inputHtml, "utf8");
// Name of resulting file
parameters["name"] = path.basename(DestinationFile);
// Set to css style margins like 10 px or 5px 5px 5px 5px.
parameters["margins"] = "5px 5px 5px 5px";
// Can be Letter, A4, A5, A6 or custom size like 200x200
parameters["paperSize"] = "Letter";
// Set to Portrait or Landscape. Portrait by default.
parameters["orientation"] = "Portrait";
// true by default. Set to false to disbale printing of background.
parameters["printBackground"] = true;
// If large input document, process in async mode by passing true
parameters["async"] = false;
// Set to HTML for header to be applied on every page at the header.
parameters["header"] = "";
// Set to HTML for footer to be applied on every page at the bottom.
parameters["footer"] = "";
// Convert JSON object to string
var jsonPayload = JSON.stringify(parameters);
// Prepare request to `HTML To PDF` API endpoint
var url = '/v1/pdf/convert/from/html';
var reqOptions = {
host: "api.pdf.co",
path: url,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
# PDF from HTML Template
Source: https://developer.pdf.co/api/pdf-from-html/convert-from-template
Convert `HTML` template into `PDF`.
## `POST /v1/pdf/convert/from/html`
Converts a predefined [HTML template](#html-templates) into a PDF document using its Template ID. The template must be created and saved in the [HTML to PDF Templates](https://app.pdf.co/html-templates-tool/manager) section of the dashboard.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`Remember to ensure that request sizes are less than `4` mb in file size. For more information, see [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `templateId` | integer | *Yes* | - | Set ID of HTML template to be used. View and manage your templates at HTML to PDF Templates. |
| `templateData` | string | *Yes* | - | Set it to a string with input `JSON` data (recommended) or `CSV` data. See [Sample JSON input](#sample-json-input) and [Sample CSV input](#sample-csv-input) for more information. |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. |
| `paperSize` | string | *No* | A4 | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. |
| `orientation` | string | *No* | Portrait | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. |
| `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL |
| `mediaType` | string | *No* | print | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). |
| `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. |
| `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. |
| `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## HTML Templates
During conversion, the API performs variable substitution and built-in helper evaluation using your [templateData](#html-templates), but does not execute any arbitrary JavaScript within the template itself. After the template is rendered into HTML, a headless browser processes the page and executes any in-page JavaScript triggered on load. External or dynamically loaded JavaScript libraries may run if accessible, but execution is not guaranteed in all environments. This separation ensures secure, predictable template handling and accurate rendering of interactive web content.
Use the dashboard to manage your HTML to [PDF Templates](https://app.pdf.co/html-templates-tool).
Templates use `{{Mustache}}` and Handlebars templating syntax. You just need to insert macros surrounded by double brackets like `{{` and `}}`.
* Find out more about [Mustache](https://mustache.github.io/mustache.5.html).
* Find out more about [Handlebars](https://handlebarsjs.com/guide/).
Some Examples of macro inside html template:
* `{{variable1}}` will be replaced with `test` if you set `templateData` to `{ "variable1": "test" }`
* `{{object1.variable1}}` will be replaced with `test` if you set `templateData` to `{ "object1": { "variable1": "test" } }`
* Simple conditions are also supported. For example: `{{#if paid}} invoice was paid {{/if}}` will show invoice was paid when `templateData` is set to `{ "paid": true }`.
### Handlebars
Handlebars extends Mustache with powerful features like conditional logic, loops, and custom helper functions. You can use Handlebars helpers like `#if`, `#unless`, `#each` (see [https://handlebarsjs.com/guide/builtin-helpers.html](https://handlebarsjs.com/guide/builtin-helpers.html)) but you can also define your own helper functions for complex calculations and data manipulation.
#### Key Handlebars Features:
* **Variable substitution**: `{{variable}}` syntax for inserting data
* **Conditional logic**: `{{#if}}` and `{{#unless}}` blocks for conditional rendering
* **Loops**: `{{#each}}` for iterating over arrays and objects
* **Nested objects**: Accessing properties with dot notation like `{{company.name}}`
* **Built-in helpers**: Using Handlebars' built-in functionality for common operations
### Sample JSON input
```json theme={null}
"templateData": "{ 'paid': true, 'invoice_id': '0002', 'total': '$999.99' }"
```
If you use `JSON` as input then make sure to escape it first (with `JSON.stringify(dataObject)` in JS). Escaping is when every `"` is replaced with `\"`. Example with `"` be escaped as `\"` then: `"templateData": "{ \"paid\": true, \"invoice_id\": \"0002\", \"total\": \"$999.99\" }"`.
### Sample CSV input
```csv theme={null}
"templateData": "paid,invoice_id,total
true,0002,$999.99"
```
## Header & Footer
The `header` and `footer` parameters can contain valid HTML markup with the following classes used to inject printing values into them:
* `date`: formatted print date
* `title`: document title
* `url`: document location
* `pageNumber`: current page number
* `totalPages`: total pages in the document
img: tag is supported in both the header and footer parameter, provided that the `src` attribute is specified as a `base64-encoded` string.
For example, the following markup will generate `Page N of NN` page numbering:
```html theme={null}
Page of .
```
### Sample Header & Footer
An example with an advanced `header` and `footer`. Note that the top and bottom page margins are important because page content may overlap the footer or header.
```json theme={null}
{
"templateId": 1,
"name": "newDocument.pdf",
"mediaType": "print",
"margins": "40px 20px 20px 20px",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "
LEFT SUBHEADERRIGHT SUBHEADER
",
"footer": "
Page of .
",
"async": false,
"templateData": "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"templateId": 1,
"name": "newDocument.pdf",
"mediaType": "print",
"margins": "40px 20px 20px 20px",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"async": false,
"templateData": "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 60646
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/html' \
--header 'x-api-key: add_your_api_key_here' \
--header 'Content-Type: application/json' \
--data-raw '{
"templateId": 2,
"name": "newDocument.pdf",
"margins": "40px 20px 20px 20px",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"async": false,
"encrypt": false,
"templateData": "{\"invoice_id\":\"1234567\",\"invoice_date\":\"April 30, 2016\",\"invoice_dateDue\":\"May 15, 2016\",\"paid\":false,\"issuer_name\":\"Acme Inc\",\"issuer_company\":\"Acme International\",\"issuer_address\":\"City, Street 3rd\",\"issuer_email\":\"support@example.com\",\"issuer_website\":\"http://example.com\",\"client_name\":\"Food Delivery Inc.\",\"client_company\":\"Food Delivery International\",\"client_address\":\"New York, Some Street, 42\",\"client_email\":\"client@example.com\",\"items\":[{\"name\":\"Setting up new web-site\",\"price\":250},{\"name\":\"Website Content Addition\",\"price\":700},{\"name\":\"Database Setup\",\"price\":200},{\"name\":\"Record Digitalization\",\"price\":1800},{\"name\":\"Cloud Storage\",\"price\":500},{\"name\":\"Short Messages\",\"price\":35},{\"name\":\"Search Engine Optimization\",\"price\":200},{\"name\":\"Priority Support\",\"price\":75},{\"name\":\"Configuring mail server and mailboxes\",\"price\":50}],\"tax\":0.065,\"discount\":0.01,\"note\":\"Thank You For Your Business!\"}"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Data to fill the template
const templateData = "./invoice_data.json";
// Destination PDF file name
const DestinationFile = "./result.pdf";
/*
Please follow below steps to create your own HTML Template and get "templateId".
1. Add new html template in app.pdf.co/templates/html
2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
3. Save this new template
4. Copy it’s ID to clipboard
5. Now set ID of the template into “templateId” parameter
*/
// HTML template using built-in template
// see https://app.pdf.co/templates/html/2/edit
const template_id = 2;
// Prepare request to `HTML To PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/html?name=${path.basename(DestinationFile)}&async=True`;
var reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json"
}
};
var requestBody = JSON.stringify({
"templateId": template_id,
"templateData": fs.readFileSync(templateData, "utf8"),
"async": true
});
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(requestBody);
postRequest.end();
function checkIfJobIsCompleted(jobId, resultFileUrl) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***********************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# --HTML Template ID--
# Please follow below steps to create your own HTML Template and get "templateId".
# 1. Add new html template in app.pdf.co/templates/html
# 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
# 3. Save this new template
# 4. Copy it’s ID to clipboard
# 5. Now set ID of the template into “templateId” parameter
# HTML template using built-in template
# see https://app.pdf.co/templates/html/2/edit
template_id = 2
# Data to fill the template
file_read = open(".\\invoice_data.json", mode='r')
TemplateData = file_read.read()
file_read.close()
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
GeneratePDFFromTemplate(template_id, TemplateData, DestinationFile)
def GeneratePDFFromTemplate(template_id, templateData, destinationFile):
"""Converts HTML to PDF using PDF.co Web API"""
data = {
'templateData': templateData,
'templateId': template_id
}
# Prepare URL for 'HTML To PDF' API request
url = "{}/pdf/convert/from/html?name={}".format(
BASE_URL,
os.path.basename(destinationFile)
)
# Execute request and get response as JSON
response = requests.post(url, data=data, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net.Http;
using System.Threading.Tasks;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoCodeSample
{
class Program
{
const String API_KEY = "***********************************";
const string DestinationFile = @".\newDocument.pdf";
const bool Async = true; // Enable asynchronous processing
static async Task Main(string[] args)
{
using (HttpClient httpClient = new HttpClient())
{
httpClient.DefaultRequestHeaders.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/pdf/convert/from/html";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary
{
{ "templateId", 1 },
{ "name", Path.GetFileName(DestinationFile) },
{ "margins", "40px 20px 20px 20px" },
{ "paperSize", "Letter" },
{ "orientation", "Portrait" },
{ "header", "" },
{ "printBackground", true },
{ "footer", "" },
{ "async", Async }, // Enable asynchronous processing
{ "encrypt", false },
{ "templateData", "{\"paid\": true,\"invoice_id\": \"0021\",\"invoice_date\": \"August 29, 2041\",\"invoice_dateDue\": \"September 29, 2041\",\"issuer_name\": \"Sarah Connor\",\"issuer_company\": \"T-800 Research Lab\",\"issuer_address\": \"435 South La Fayette Park Place, Los Angeles, CA 90057\",\"issuer_website\": \"www.example.com\",\"issuer_email\": \"info@example.com\",\"client_name\": \"Cyberdyne Systems\",\"client_company\": \"Cyberdyne Systems\",\"client_address\": \"18144 El Camino Real, Sunnyvale, California\",\"client_email\": \"sales@example.com\",\"items\": [ { \"name\": \"T-800 Prototype Research\", \"price\": 1000.00 }, { \"name\": \"T-800 Cloud Sync Setup\", \"price\": 300.00 } ],\"discount\": 100,\"tax\": 87,\"total\": 1287,\"note\": \"Thank you for your support of advanced robotics.\"}" }
};
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Send POST request asynchronously
var content = new StringContent(jsonPayload, System.Text.Encoding.UTF8, "application/json");
HttpResponseMessage response = await httpClient.PostAsync(url, content);
// Read response asynchronously
string responseBody = await response.Content.ReadAsStringAsync();
JObject json = JObject.Parse(responseBody);
if (json["error"].ToObject() == false)
{
// Get Job ID for asynchronous processing
string jobId = json["jobId"].ToString();
Console.WriteLine($"Job ID: {jobId}");
// Check job status in a loop
bool isJobCompleted = false;
while (!isJobCompleted)
{
string status = await CheckJobStatusAsync(httpClient, jobId);
Console.WriteLine($"Job Status: {status}");
if (status == "success")
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file asynchronously
byte[] fileBytes = await httpClient.GetByteArrayAsync(resultFileUrl);
await File.WriteAllBytesAsync(DestinationFile, fileBytes);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
isJobCompleted = true;
}
else if (status == "working")
{
// Wait for a few seconds before checking again
await Task.Delay(3000);
}
else
{
Console.WriteLine($"Job failed or aborted. Status: {status}");
isJobCompleted = true;
}
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (Exception e)
{
Console.WriteLine(e.ToString());
}
}
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
static async Task CheckJobStatusAsync(HttpClient httpClient, string jobId)
{
string url = $"https://api.pdf.co/v1/job/check?jobid={jobId}";
// Send GET request to check job status
HttpResponseMessage response = await httpClient.GetAsync(url);
string responseBody = await response.Content.ReadAsStringAsync();
JObject json = JObject.Parse(responseBody);
return json["status"].ToString();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
/*
Please follow below steps to create your own HTML Template and get "templateId".
1. Add new html template in app.pdf.co/templates/html
2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
3. Save this new template
4. Copy it’s ID to clipboard
5. Now set ID of the template into “templateId” parameter
*/
// HTML template using built-in template
// see https://app.pdf.co/templates/html/2/edit
final String templateId = "2";
// Data to fill the template
final String templateData = new String(Files.readAllBytes(Paths.get(".\\invoice_data.json")));
// Destination PDF file name
final Path destinationFile = Paths.get(".\\result.pdf");
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `HTML to PDF` API call
String query = String.format(
"https://api.pdf.co/v1/pdf/convert/from/html?name=%s",
destinationFile.getFileName());
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Prepare request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("templateId", new JsonPrimitive(templateId));
jsonBody.add("templateData", new JsonPrimitive(templateData));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF Invoice Generation Results
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
```
# Return HTML Template by ID
Source: https://developer.pdf.co/api/pdf-from-html/template-id
Returns HTML template by template’s id.
**Try it live:** [Return HTML Template by ID → API Tester](/api-tester/pdf-from-html/template-id) — send a real request from your browser.
## `GET /templates/html/:id`
Once you have obtained a template then use the [PDF from HTML](/api/pdf-from-html) API with the required `templateId` & `templateData` parameters defined.
Use the dashboard to manage your [HTML to PDF Templates](https://app.pdf.co/html-templates-tool).
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------ |
| `id` | integer | Template ID. |
| `type` | string | Template type. |
| `title` | string | Template title. |
| `description` | string | Template description. |
| `test_json` | string | Template test JSON. |
| `updated_at` | string | Template updated at. |
| `body` | string | Template content. |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `credits` | integer | Number of credits consumed by the request |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"templates": [
{
"id": 1,
"type": "system",
"title": "General Invoice Template",
"description": "sample invoice template showcasing use of Mustache templates syntax for generating invoices"
},
{
"id": 15,
"type": "user",
"title": "User Template 1",
"description": ""
}
],
"remainingCredits": 99204004,
"credits": 2
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/templates/html' \
--header 'Content-Type: application/json' \
--header 'x-api-key: '
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Data to fill the template
const templateData = "./invoice_data.json";
// Destination PDF file name
const DestinationFile = "./result.pdf";
/*
Please follow below steps to create your own HTML Template and get "templateId".
1. Add new html template in app.pdf.co/templates/html
2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
3. Save this new template
4. Copy it’s ID to clipboard
5. Now set ID of the template into “templateId” parameter
*/
// HTML template using built-in template
// see https://app.pdf.co/templates/html/2/edit
const template_id = 2;
// Prepare request to `HTML To PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/html?name=${path.basename(DestinationFile)}&async=True`;
var reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json"
}
};
var requestBody = JSON.stringify({
"templateId": template_id,
"templateData": fs.readFileSync(templateData, "utf8"),
"async": true
});
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(requestBody);
postRequest.end();
function checkIfJobIsCompleted(jobId, resultFileUrl) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function(){ checkIfJobIsCompleted(jobId, resultFileUrl);}, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***********************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# --HTML Template ID--
# Please follow below steps to create your own HTML Template and get "templateId".
# 1. Add new html template in app.pdf.co/templates/html
# 2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
# 3. Save this new template
# 4. Copy it’s ID to clipboard
# 5. Now set ID of the template into “templateId” parameter
# HTML template using built-in template
# see https://app.pdf.co/templates/html/2/edit
template_id = 2
# Data to fill the template
file_read = open(".\\invoice_data.json", mode='r')
TemplateData = file_read.read()
file_read.close()
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
GeneratePDFFromTemplate(template_id, TemplateData, DestinationFile)
def GeneratePDFFromTemplate(template_id, templateData, destinationFile):
"""Converts HTML to PDF using PDF.co Web API"""
data = {
'templateData': templateData,
'templateId': template_id
}
# Prepare URL for 'HTML To PDF' API request
url = "{}/pdf/convert/from/html?name={}".format(
BASE_URL,
os.path.basename(destinationFile)
)
# Execute request and get response as JSON
response = requests.post(url, data=data, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
static void Main(string[] args)
{
// --TemplateID--
/*
Please follow below steps to create your own HTML Template and get "templateId".
1. Add new html template in app.pdf.co/templates/html
2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
3. Save this new template
4. Copy it’s ID to clipboard
5. Now set ID of the template into “templateId” parameter
*/
// HTML template using built-in template
// see https://app.pdf.co/templates/html/2/edit
var templateId = 2;
// Data to fill the template
string templateData = File.ReadAllText(@".\invoice_data.json");
// Destination PDF file name
string destinationFile = @".\result.pdf";
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
webClient.Headers.Add("Content-Type", "application/json");
try
{
// URL for `HTML to PDF` API call
string url = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/pdf/convert/from/html?name={0}",
Path.GetFileName(destinationFile)));
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(destinationFile));
parameters.Add("templateId", templateId);
parameters.Add("templateData", templateData);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute request
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
webClient.Headers.Remove("Content-Type"); // remove the header required for only the previous request
// Download the PDF file
webClient.DownloadFile(resultFileUrl, destinationFile);
Console.WriteLine("Generated PDF document saved as \"{0}\" file.", destinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key to exit...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import com.google.gson.JsonPrimitive;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
public static void main(String[] args) throws IOException
{
/*
Please follow below steps to create your own HTML Template and get "templateId".
1. Add new html template in app.pdf.co/templates/html
2. Copy paste your html template code into this new template. Sample HTML templates can be found at "https://github.com/pdfdotco/pdf-co-api-samples/tree/master/PDF%20from%20HTML%20template/TEMPLATES-SAMPLES"
3. Save this new template
4. Copy it’s ID to clipboard
5. Now set ID of the template into “templateId” parameter
*/
// HTML template using built-in template
// see https://app.pdf.co/templates/html/2/edit
final String templateId = "2";
// Data to fill the template
final String templateData = new String(Files.readAllBytes(Paths.get(".\\invoice_data.json")));
// Destination PDF file name
final Path destinationFile = Paths.get(".\\result.pdf");
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `HTML to PDF` API call
String query = String.format(
"https://api.pdf.co/v1/pdf/convert/from/html?name=%s",
destinationFile.getFileName());
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Prepare request body in JSON format
JsonObject jsonBody = new JsonObject();
jsonBody.add("templateId", new JsonPrimitive(templateId));
jsonBody.add("templateData", new JsonPrimitive(templateData));
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonBody.toString());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF Invoice Generation Results
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
```
# Return All Templates
Source: https://developer.pdf.co/api/pdf-from-html/templates
Create PDF’s from HTML template input.
**Try it live:** [Return All Templates → API Tester](/api-tester/pdf-from-html/templates) — send a real request from your browser.
## `GET /templates/html`
Once you have obtained a template then use the [PDF from HTML](/api/pdf-from-html) API with the required templateId & templateData parameters defined.
Use the dashboard to manage your [HTML to PDF Templates](https://app.pdf.co/html-templates-tool).
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | -------------- | ------------------------------------------ |
| `templates` | array\[object] | |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `credits` | integer | Number of credits consumed by the request |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"id": 1,
"type": "system",
"title": "General Invoice Template",
"description": "sample invoice template showcasing use of Mustache templates syntax for generating invoices",
"body": "\r\n\r\n\r\nInvoice \r\n\r\n \r\n\r\n \r\n
\r\n PAID
\r\n \r\n \r\n
\r\n
\r\n
\r\n \r\n \r\n
\r\n
\r\n \r\n\r\n \r\n \r\n \r\n \r\n
\r\n
\r\n
\r\n
\r\n Invoice Number: \r\n
\r\n
\r\n Invoice Date: \r\n
\r\n
\r\n Invoice Due Date: \r\n
\r\n
\r\n
\r\n
\r\n \r\n
\r\n \r\n\r\n
\r\n
BILL TO
\r\n
\r\n
Name:
\r\n
Company:
\r\n
Address:
\r\n
Email:
\r\n
\r\n
\r\n
\r\n \r\n
\r\n
\r\n
\r\n \r\n
\r\n
Item
\r\n
Price
\r\n
\r\n \r\n \r\n \r\n
\r\n
\r\n
\r\n
\r\n \r\n \r\n
\r\n
\r\n\r\n
\r\n
\r\n
\r\n
\r\n
\r\n
Discount:
\r\n
Tax:
\r\n
TOTAL:
\r\n
\r\n \r\n
\r\n
\r\n
\r\n
\r\n \r\n \r\n\r\n\r\n",
"remainingCredits": 99204002,
"credits": 2
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request GET 'https://api.pdf.co/v1/templates/html/1' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw ''
```
# PDF from Image
Source: https://developer.pdf.co/api/pdf-from-image
Convert `JPG`, `PNG`, `TIFF` image formats into `PDF`.
**Try it live:** [PDF from Image → API Tester](/api-tester/pdf-from-image) — send a real request from your browser.
## `POST /pdf/convert/from/image`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URLs to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources). If you use multiple URLs, please separate them with a `,` |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
To add multiple images to convert pleaase comma-separate the `url` parameter. e.g. "url": `"https://example.com/image1.png,https://example.com/image2.png"`
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png,https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/5ef3d4033e344ec091bbeb3cfd848633/image2.pdf",
"pageCount": 2,
"error": false,
"status": 200,
"name": "image2.pdf",
"remainingCredits": 59871
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/image' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png,https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URLs of image files to convert to PDF document
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFiles = [
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg"
];
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Image To PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/image`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceFiles.join(",")
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URLs of image files to convert to PDF document
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFiles = [
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg"
]
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
SourceFileURL = ",".join(SourceFiles)
convertImageToPDF(SourceFileURL, DestinationFile)
def convertImageToPDF(uploadedFileUrl, destinationFile):
"""Converts Image to PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrl
# Prepare URL for 'Image To PDF' API request
url = "{}/pdf/convert/from/image".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URLs of image files to convert to PDF document
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
static string[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" };
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Image To PDF` API call
string url = "https://api.pdf.co/v1/pdf/convert/from/image";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", string.Join(",", SourceFiles));
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URLs of image files to convert to PDF document
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String[] SourceFiles = {
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png",
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image2.jpg" };
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Image To PDF` API call
String query = "https://api.pdf.co/v1/pdf/convert/from/image";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
String.join(",", SourceFiles));
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Image To PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF from URL
Source: https://developer.pdf.co/api/pdf-from-url
Convert any web page URL into a PDF document with custom margins, paper size, orientation, headers and footers.
**Try it live:** [PDF from URL → API Tester](/api-tester/pdf-from-url) — send a real request from your browser.
## `POST /v1/pdf/convert/from/url`
This method will process any JavaScript which the webpage triggers when it loads. For example if the the webpage triggers a JavaScript popup window then that will be included in the conversion process. There is no option to disable JavaScript on the supplied `HTML` page.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `margins` | string | *No* | - | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. |
| `paperSize` | string | *No* | `A4` | Specifies the paper size. Accepts standard sizes like 'Letter', 'Legal', 'Tabloid', 'Ledger', 'A0'–'A6'. You can also set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: '200 300', '200px 300px', '200mm 300mm', '20cm 30cm', '6in 8in'. |
| `orientation` | string | *No* | `Portrait` | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. |
| `printBackground` | boolean | *No* | `true` | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL |
| `mediaType` | string | *No* | `print` | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). |
| `DoNotWaitFullLoad` | boolean | *No* | `false` | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- false waits for full page load, while true speeds up conversion by waiting only for minimal loading. |
| `renderTimeout` | integer | *No* | - | Specifies the maximum time (in milliseconds) to wait for the page to fully render before generating the PDF. Useful for pages with heavy JavaScript, dynamic content, or slow-loading resources. The maximum allowed value is 177000ms (177 seconds). |
| `legacyRendition` | boolean | *No* | `false` | Forces rendering with an older Chromium version (`r970485`) to work around version-specific issues where modern browsers can't paginate certain page structures correctly. Ensures full content is rendered as a long, printable page. |
| `header` | string | *No* | - | Set this to can add user definable HTML for the header to be applied on every page header. The format is html. |
| `footer` | string | *No* | - | Set this to can add user definable HTML for the footer to be applied on every page bottom. The format is html. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Header & Footer
The header and footer parameters can contain valid HTML markup with the following classes used to inject printing values into them:
date: formatted print date
title: document title
url: document location
pageNumber: current page number
totalPages: total pages in the document
img: tag is supported in both the header and footer parameter, provided that the src attribute is specified as a base64-encoded string.
For example, the following markup will generate Page N of NN page numbering:
```html theme={null}
Page of .
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://wikipedia.org/wiki/Wikipedia:Contact_us",
"name": "result.pdf",
"margins": "5mm",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"mediaType": "print",
"renderTimeout": 15000,
"async": false,
"profiles": "{ \"CustomScript\": \";; // put some custom js script here \"}"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/97dc323f32794eae8fa6602f5bd981c1/result.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 60646
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/from/url' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://wikipedia.org/wiki/Wikipedia:Contact_us",
"margins": "5mm",
"paperSize": "Letter",
"orientation": "Portrait",
"printBackground": true,
"header": "",
"footer": "",
"mediaType": "print",
"async": false,
"profiles": "{ \"CustomScript\": \";; // put some custom js script here \"}"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// URL of web page to convert to PDF document.
const SourceUrl = "http://en.wikipedia.org/wiki/Main_Page";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Web Page to PDF` API endpoint
var queryPath = `/v1/pdf/convert/from/url`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), url: SourceUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "**********************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# URL of web page to convert to PDF document.
SourceUrl = "http://en.wikipedia.org/wiki/Main_Page"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
convertHTMLToPDF(SourceUrl, DestinationFile)
def convertHTMLToPDF(uploadedFileUrl, destinationFile):
"""Converts HTML to PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["url"] = uploadedFileUrl
# Prepare URL for 'HTML To PDF' API request
url = "{}/pdf/convert/from/url".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// URL of web page to convert to PDF document.
const string SourceUrl = "http://en.wikipedia.org/wiki/Main_Page";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Web Page to PDF` API call
string url = "https://api.pdf.co/v1/pdf/convert/from/url";
// Prepare requests params as JSON
Dictionary requestBody = new Dictionary();
requestBody.Add("name", Path.GetFileName(DestinationFile));
requestBody.Add("url", SourceUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(requestBody);
try
{
// Execute POST request
var response = webClient.UploadString(url, "POST", jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF document saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// URL of web page to convert to PDF document.
final static String SourceUrl = "http://en.wikipedia.org/wiki/Main_Page";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Web Page to PDF` API call
String query = "https://api.pdf.co/v1/pdf/convert/from/url";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
SourceUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF Extractor Results
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
?>
```
# PDF Info Reader
Source: https://developer.pdf.co/api/pdf-info-reader
Get detailed information about a **PDF** document, it’s properties and security permissions.
**Try it live:** [PDF Info Reader → API Tester](/api-tester/pdf-info-reader) — send a real request from your browser.
## `POST /v1/pdf/info`
Extracts basic information about an input PDF file, PDF file security permissions, and other information. If you want to extract information about fillable fields (checkboxes, radiobuttons, listboxes) from PDF then please use [/pdf/info/fields](/api/forms/info-reader) instead.
For one-time check of PDF file information and find form fields please use PDF [Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper).
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters.AddGammaCorrection()` | array\[string (float format)] | *No* | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. |
| `OCRImagePreprocessingFilters.AddGrayscale()` | boolean | *No* | `false` | Set to true to preprocessing filter that converts a colored document/image to grayscale before performing OCR |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `info` | object | Info details. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"info": {
"PageCount": 1,
"Author": "Alice V. Knox",
"Title": "Kid's News 1",
"Producer": "Acrobat Distiller 4.0 for Windows",
"Subject": "Kid's News 1",
"CreationDate": "8/15/2001 2:50:36 PM",
"Bookmarks": "",
"Keywords": "",
"Creator": "Adobe PageMaker 6.52",
"Encrypted": false,
"PageRectangle": {
"Location": {
"IsEmpty": true,
"X": 0,
"Y": 0
},
"Size": "612, 792",
"X": 0,
"Y": 0,
"Width": 612,
"Height": 792,
"Left": 0,
"Top": 0,
"Right": 612,
"Bottom": 792,
"IsEmpty": false
},
"ModificationDate": "9/20/2001 6:23:02 PM",
"EncryptionAlgorithm": 0,
"PermissionPrinting": true,
"PermissionModifyDocument": true,
"PermissionContentExtraction": true,
"PermissionModifyAnnotations": true,
"PermissionFillForms": true,
"PermissionAccessibility": true,
"PermissionAssemble": true,
"PermissionHighQualityPrint": true
},
"error": false,
"status": 200,
"remainingCredits": 77732
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/info' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-info/sample.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file to get information
const SourceFile = "./sample.pdf";
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. GET INFORMATION FROM UPLOADED FILE
getPdfInfo(API_KEY, uploadedFileUrl);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function getPdfInfo(apiKey, uploadedFileUrl) {
// Prepare URL for `PDF Info` API call
var queryPath = `/v1/pdf/info`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Display PDF document information
for (var key in data.info) {
console.log(`${key}: ${data.info[key]}`);
}
}
else {
// Service reported error
console.log("getPdfInfo(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPdfInfo(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
getInfoFromPDF(uploadedFileUrl)
def getInfoFromPDF(uploadedFileUrl):
"""Get Information using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF Info' API request
url = "{}/pdf/info".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Display information
print(json["info"])
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file to get information
const string SourceFile = @".\sample.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. GET INFORMATION FROM UPLOADED FILE
// URL for `PDF Info` API call
var url = "https://api.pdf.co/v1/pdf/info";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Display PDF document information
foreach (JToken token in json["info"])
{
JProperty property = (JProperty) token;
Console.WriteLine("{0}: {1}", property.Name, property.Value);
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.util.Map;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source file name
final static Path SourceFile = Paths.get(".\\sample.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, uploadUrl, SourceFile.toFile()))
{
// 3. GET INFORMATION FROM UPLOADED FILE
getPdfInfo(webClient, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void getPdfInfo(OkHttpClient webClient, String uploadedFileUrl) throws IOException {
// Prepare URL for `PDF Info` API call
String query = "https://api.pdf.co/v1/pdf/info";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"url\": \"%s\"}",
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Display PDF document information
JsonObject info = (JsonObject) json.get("info");
for (Map.Entry entry : info.entrySet())
{
System.out.println(entry.getKey() + ": " + entry.getValue());
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String url, File sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
}
```
```php theme={null}
PDF Information Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
}
?>
```
# Add Password to PDF
Source: https://developer.pdf.co/api/pdf-password/add
Add password and security limitations to PDF
**Try it live:** [Add Password to PDF → API Tester](/api-tester/pdf-password/add) — send a real request from your browser.
## `POST /v1/pdf/security/add`
**Modifying Restriction Settings**: To modify assembly or extraction settings that control PDF restrictions (such as `allowPrintDocument`, `allowFillForms`, `allowModifyDocument`, `allowAssemblyDocument`, and related permission settings), you **must** use the `ownerPassword` in your request. The `userPassword` alone cannot modify these permission settings. Attempting to change these restrictions with only a `userPassword` will result in an error: `"This file is password-protected. Please ensure you've entered the correct password…"`. This requirement applies only when making changes to those restrictions.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `ownerPassword` | string | *No* | - | The main owner password that is used for document encryption and for setting/removing restrictions. |
| `userPassword` | string | *No* | - | The optional user password will be asked for viewing and printing document. |
| `encryptionAlgorithm` | string | *No* | AES\_128bit | Encryption algorithm. AES\_128bit or higher is recommended. The available algorithms are: `RC4_40bit`, `RC4_128bit`, `AES_128bit`, `AES_256bit`. |
| `allowAccessibilitySupport` | boolean | *No* | `false` | Allow or prohibit content extraction for accessibility needs. |
| `allowAssemblyDocument` | boolean | *No* | `false` | Allow or prohibit assembling the document. |
| `allowPrintDocument` | boolean | *No* | `false` | Allow or prohibit printing PDF document. |
| `allowFillForms` | boolean | *No* | `false` | Allow or prohibit the filling of interactive form fields (including signature fields) in the PDF documents. |
| `allowModifyDocument` | boolean | *No* | `false` | Allow or prohibit modification of PDF document. |
| `allowContentExtraction` | boolean | *No* | `false` | Allow or prohibit copying content from PDF document. |
| `allowModifyAnnotations` | boolean | *No* | `false` | Allow or prohibit interacting with text annotations and forms in PDF document. |
| `printQuality` | string | *No* | HighResolution | Allowed printing quality. The available modes are: `LowResolution`, `HighResolution`. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
This restriction applies when `userPassword` (if any) is entered. This restriction does not apply if the user enters `ownerPassword`.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"ownerPassword": "12345",
"userPassword": "54321",
"EncryptionAlgorithm": "AES_128bit",
"AllowPrintDocument": false,
"AllowFillForms": false,
"AllowModifyDocument": false,
"AllowContentExtraction": false,
"AllowModifyAnnotations": false,
"PrintQuality": "LowResolution",
"name": "output-protected.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/eaa441ade38548b8a3a96d8014c4f463/sample1.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample1.pdf",
"remainingCredits": 616208,
"credits": 14
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/security/add' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf",
"ownerPassword": "12345",
"userPassword": "54321",
"EncryptionAlgorithm": "AES_128bit",
"AllowPrintDocument": false,
"AllowFillForms": false,
"AllowModifyDocument": false,
"AllowContentExtraction": false,
"AllowModifyAnnotations": false,
"PrintQuality": "LowResolution",
"name": "output-protected.pdf",
"async": false
}'
```
# Remove Password from PDF
Source: https://developer.pdf.co/api/pdf-password/remove
Remove existing limits and password from PDF file.
**Try it live:** [Remove Password from PDF → API Tester](/api-tester/pdf-password/remove) — send a real request from your browser.
## `POST /v1/pdf/security/remove`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-security/ProtectedPDFFile.pdf",
"password": "admin@123",
"name": "unprotected",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/9f2a754f76db46ac93781b3d2c6694c3/ProtectedPDFFile.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "ProtectedPDFFile.pdf",
"remainingCredits": 616187,
"credits": 21
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/security/remove' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-security/ProtectedPDFFile.pdf",
"password": "admin@123",
"name": "unprotected",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf";
// Destination PDF file name
const DestinationFile = "./protected.pdf";
// Passwords to protect PDF document
// The owner password will be required for document modification.
// The user password only allows to view and print the document.
const OwnerPassword = "123456";
const UserPassword = "654321";
// Encryption algorithm.
// Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit".
const EncryptionAlgorithm = "AES_128bit";
// Allow or prohibit content extraction for accessibility needs.
const AllowAccessibilitySupport = true;
// Allow or prohibit assembling the document.
const AllowAssemblyDocument = true;
// Allow or prohibit printing PDF document.
const AllowPrintDocument = true;
// Allow or prohibit filling of interactive form fields (including signature fields) in PDF document.
const AllowFillForms = true;
// Allow or prohibit modification of PDF document.
const AllowModifyDocument = true;
// Allow or prohibit copying content from PDF document.
const AllowContentExtraction = true;
// Allow or prohibit interacting with text annotations and forms in PDF document.
const AllowModifyAnnotations = true;
// Allowed printing quality.
// Valid values: "HighResolution", "LowResolution"
const PrintQuality = "HighResolution";
// Runs processing asynchronously.
// Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success).
const async = false;
// Prepare request to `PDF Security` API endpoint
var queryPath = `/v1/pdf/security/add`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile),
url: SourceFileUrl,
ownerPassword: OwnerPassword,
userPassword: UserPassword,
encryptionAlgorithm: EncryptionAlgorithm,
allowAccessibilitySupport: AllowAccessibilitySupport,
allowAssemblyDocument: AllowAssemblyDocument,
allowPrintDocument: AllowPrintDocument,
allowFillForms: AllowFillForms,
allowModifyDocument: AllowModifyDocument,
allowContentExtraction: AllowContentExtraction,
allowModifyAnnotations: AllowModifyAnnotations,
printQuality: PrintQuality,
async: async
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "********************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf"
# Destination PDF file name
DestinationFile = ".\\protected.pdf"
# Passwords to protect PDF document
# The owner password will be required for document modification.
# The user password only allows to view and print the document.
OwnerPassword = "123456"
UserPassword = "654321"
# Encryption algorithm.
# Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit".
EncryptionAlgorithm = "AES_128bit"
# Allow or prohibit content extraction for accessibility needs.
AllowAccessibilitySupport = True
# Allow or prohibit assembling the document.
AllowAssemblyDocument = True
# Allow or prohibit printing PDF document.
AllowPrintDocument = True
# Allow or prohibit filling of interactive form fields (including signature fields) in PDF document.
AllowFillForms = True
# Allow or prohibit modification of PDF document.
AllowModifyDocument = True
# Allow or prohibit copying content from PDF document.
AllowContentExtraction = True
# Allow or prohibit interacting with text annotations and forms in PDF document.
AllowModifyAnnotations = True
# Allowed printing quality.
# Valid values: "HighResolution", "LowResolution"
PrintQuality = "HighResolution"
# Runs processing asynchronously.
# Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success).
Async = False
def main(args = None):
protectPDF(SourceFileURL, DestinationFile)
def protectPDF(uploadedFileUrl, destinationFile):
"""Protect PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {"name": os.path.basename(destinationFile), "url": uploadedFileUrl, "ownerPassword": OwnerPassword,
"userPassword": UserPassword, "encryptionAlgorithm": EncryptionAlgorithm,
"allowAccessibilitySupport": AllowAccessibilitySupport,
"allowAssemblyDocument": AllowAssemblyDocument, "allowPrintDocument": AllowPrintDocument,
"allowFillForms": AllowFillForms, "allowModifyDocument": AllowModifyDocument,
"allowContentExtraction": AllowContentExtraction, "allowModifyAnnotations": AllowModifyAnnotations,
"printQuality": PrintQuality, "async": Async}
# Serializing json
import json
json_object = json.dumps(parameters, indent=4)
# Prepare URL for 'PDF Security' API request
url = "{}/pdf/security/add".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=json_object, headers={"x-api-key": API_KEY})
if (response.status_code == 200):
jsonResp = response.json()
if jsonResp["error"] == False:
# Get URL of result file
resultFileUrl = jsonResp["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(jsonResp["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.CodeDom;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Destination PDF file name
const string DestinationFile = @".\protected.pdf";
// Passwords to protect PDF document
// The owner password will be required for document modification.
// The user password only allows to view and print the document.
const string OwnerPassword = "123456";
const string UserPassword = "654321";
// Encryption algorithm.
// Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit".
const string EncryptionAlgorithm = "AES_128bit";
// Allow or prohibit content extraction for accessibility needs.
const bool AllowAccessibilitySupport = true;
// Allow or prohibit assembling the document.
const bool AllowAssemblyDocument = true;
// Allow or prohibit printing PDF document.
const bool AllowPrintDocument = true;
// Allow or prohibit filling of interactive form fields (including signature fields) in PDF document.
const bool AllowFillForms = true;
// Allow or prohibit modification of PDF document.
const bool AllowModifyDocument = true;
// Allow or prohibit copying content from PDF document.
const bool AllowContentExtraction = true;
// Allow or prohibit interacting with text annotations and forms in PDF document.
const bool AllowModifyAnnotations = true;
// Allowed printing quality.
// Valid values: "HighResolution", "LowResolution"
const string PrintQuality = "HighResolution";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Upload file to the cloud
string uploadedFileUrl = UploadFile(SourceFile);
// PROTECT UPLOADED PDF DOCUMENT
// Prepare requests params as JSON
// See documentation: https://developer.pdf.co/
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("url", uploadedFileUrl);
parameters.Add("ownerPassword", OwnerPassword);
parameters.Add("userPassword", UserPassword);
parameters.Add("encryptionAlgorithm", EncryptionAlgorithm);
parameters.Add("allowAccessibilitySupport", AllowAccessibilitySupport.ToString());
parameters.Add("allowAssemblyDocument", AllowAssemblyDocument.ToString());
parameters.Add("allowPrintDocument", AllowPrintDocument.ToString());
parameters.Add("allowFillForms", AllowFillForms.ToString());
parameters.Add("allowModifyDocument", AllowModifyDocument.ToString());
parameters.Add("allowContentExtraction", AllowContentExtraction.ToString());
parameters.Add("allowModifyAnnotations", AllowModifyAnnotations.ToString());
parameters.Add("printQuality", PrintQuality);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// URL of "PDF Security" endpoint
string url = "https://api.pdf.co/v1/pdf/security/add";
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download generated PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
///
/// Uploads file to the cloud and return URL of uploaded file to use in further API calls.
///
/// Source file name (path).
/// URL of uploaded file
static string UploadFile(string file)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
try
{
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(file)));
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
// Get URL of uploaded file to use with later API calls
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", file); // You can use UploadData() instead if your file is in byte[] or Stream
return uploadedFileUrl;
}
else
{
// Display service reported error
Console.WriteLine(json["message"].ToString());
}
}
catch (Exception e)
{
Console.WriteLine(e);
throw;
}
finally
{
webClient.Dispose();
}
return null;
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-merge/sample1.pdf";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\protected.pdf");
// Passwords to protect PDF document
// The owner password will be required for document modification.
// The user password only allows to view and print the document.
final static String OwnerPassword = "123456";
final static String UserPassword = "654321";
// Encryption algorithm.
// Valid values: "RC4_40bit", "RC4_128bit", "AES_128bit", "AES_256bit".
final static String EncryptionAlgorithm = "AES_128bit";
// Allow or prohibit content extraction for accessibility needs.
final static boolean AllowAccessibilitySupport = true;
// Allow or prohibit assembling the document.
final static boolean AllowAssemblyDocument = true;
// Allow or prohibit printing PDF document.
final static boolean AllowPrintDocument = true;
// Allow or prohibit filling of interactive form fields (including signature fields) in PDF document.
final static boolean AllowFillForms = true;
// Allow or prohibit modification of PDF document.
final static boolean AllowModifyDocument = true;
// Allow or prohibit copying content from PDF document.
final static boolean AllowContentExtraction = true;
// Allow or prohibit interacting with text annotations and forms in PDF document.
final static boolean AllowModifyAnnotations = true;
// Allowed printing quality.
// Valid values: "HighResolution", "LowResolution"
final static String PrintQuality = "HighResolution";
// Runs processing asynchronously.
// Returns Use JobId that you may use with /job/check to check state of the processing (possible states: working, failed, aborted and success).
final static boolean async = false;
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `PDF Security` API call
String query = "https://api.pdf.co/v1/pdf/security/add";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\n" +
" \"name\": \"%s\",\n" +
" \"url\": \"%s\",\n" +
" \"ownerPassword\": \"%s\",\n" +
" \"userPassword\": \"%s\",\n" +
" \"EncryptionAlgorithm\": \"%s\",\n" +
" \"AllowAccessibilitySupport\": %s,\n" +
" \"AllowAssemblyDocument\": %s,\n" +
" \"AllowPrintDocument\": %s,\n" +
" \"AllowFillForms\": %s,\n" +
" \"AllowModifyDocument\": %s,\n" +
" \"AllowContentExtraction\": %s,\n" +
" \"AllowModifyAnnotations\": %s,\n" +
" \"PrintQuality\": \"%s\",\n" +
" \"async\": %s\n" +
"}",
DestinationFile.getFileName(), SourceFileUrl, OwnerPassword, UserPassword, EncryptionAlgorithm,
Boolean.toString(AllowAccessibilitySupport), Boolean.toString(AllowAssemblyDocument), Boolean.toString(AllowPrintDocument),
Boolean.toString(AllowFillForms), Boolean.toString(AllowModifyDocument), Boolean.toString(AllowContentExtraction),
Boolean.toString(AllowModifyAnnotations),PrintQuality, Boolean.toString(async)
);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# Auto-rotate Pages with AI
Source: https://developer.pdf.co/api/pdf-rotate/auto
Automatically detect and fix page rotation in scanned PDFs using AI-based text analysis, with a configurable OCR language parameter.
**Try it live:** [Auto-rotate Pages with AI → API Tester](/api-tester/pdf-rotate/auto) — send a real request from your browser.
## `POST /v1/pdf/edit/rotate/auto`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `fileSize` | integer | Size of the optimized PDF file in bytes |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf",
"name": "result.pdf"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/HQ86WA7MFED1Q7C843NDSKVRW5AFVTMP/result.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEFwaDD6ndEhlfId4KouQ8yKCASr8amowIV2tLAi%2BhjnlVi%2FNYjf8ZJ3MqgKWsYVm5dQ8fQx7hceGdmtqhB6OH8t9xdacbMEcoMIpQr1BcSSfu2ZFfGFBDaHNpaSTfPXhkNnQaZFOi5KFozZiPBP9xPoSCV%2Fj%2BLIrDsOF%2Fb89i1Nd4OJFoXnfhjf03ZHJ%2BNCQEC%2BbePsovsjNlQYyKIV1qb1wGdwgJ%2ByibJ5x%2BQGoG4x2ebnEGTQkKBf4zobYT9Uv6FVQ%2FJg%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHB3D5EZ56/20220623/us-west-2/s3/aws4_request&X-Amz-Date=20220623T141005Z&X-Amz-SignedHeaders=host&X-Amz-Signature=351ad45980cad0d3f634f7b784d5445b9c4e1ad2244912f56ea065209d73bb30",
"fileSize": 455115,
"pageCount": 3,
"error": false,
"status": 200,
"name": "result.pdf",
"credits": 84,
"duration": 6002,
"remainingCredits": 98194629
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/rotate/auto' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf",
"name": "result.pdf"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Rotate PDF` API endpoint
var queryPath = `/v1/pdf/edit/rotate/auto`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
url: SourceFileUrl, name: path.basename(DestinationFile), angle: Angle, pages: Pages
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
rotatePDF(SourceFileURL, DestinationFile)
def rotatePDF(uploadedFileUrl, destinationFile):
"""Auto Rotate PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["url"] = uploadedFileUrl
parameters["name"] = os.path.basename(destinationFile)
# Prepare URL for 'Auto Rotate PDF' API request
url = "{}/pdf/edit/rotate/auto".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Auto Rotate PDF` API call
string url = "https://api.pdf.co/v1/pdf/edit/rotate/auto";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("url", SourceFileUrl);
parameters.Add("name", Path.GetFileName(DestinationFile));
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Rotate PDF` API call
String query = "https://api.pdf.co/v1/pdf/edit/rotate/auto";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"url\": \"%s\", \"name\": \"%s\"}",
SourceFileUrl,
DestinationFile.getFileName());
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Optimize PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Rotate Selected Pages
Source: https://developer.pdf.co/api/pdf-rotate/basic
Rotates selected pages inside a PDF file.
**Try it live:** [Rotate Selected Pages → API Tester](/api-tester/pdf-rotate/basic) — send a real request from your browser.
## `POST /v1/pdf/edit/rotate`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `angle` | integer | *No* | `0` | Contains the angle of rotation for the PDF pages. The available angles are: `0`, `90`, `180`, `270`. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `fileSize` | integer | Size of the optimized PDF file in bytes |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-fix-rotation/rotated_pages.pdf",
"name": "result.pdf"
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/2EK4QYIZU1XUEUH853VTSK47NPLXCUYX/result.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEFsaDC5Vfgoi83YzdW9HXiKCAYBVHK096wqoUyu8Ckq8jEhV1DBv9VzHY1EcPWvfG3L2YrFa8QC5ZMr3UhFEn4%2B%2B2u6e%2FcdZd%2FXbdVaI45yNE%2Btz28UHMVxCQUClj9kCHrMyJ4W1%2BlnDgLi9JUHt7SkIvV9Lj7GLDBOXy22KCND86HdtPg0uT%2FNQtcjJm%2F34cISImKYov63NlQYyKG%2BEO%2FQLP%2BzJuugBdSLcKUOTL52dnc1l82ye1u5kYvTlbfPMdisU1tY%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHESTCUXXF/20220623/us-west-2/s3/aws4_request&X-Amz-Date=20220623T135902Z&X-Amz-SignedHeaders=host&X-Amz-Signature=78f2fe7f22ebd1fc2c329a855c7582f9deb09c0c5045281b9b2420a5afb792cf",
"fileSize": 1064923,
"pageCount": 4,
"error": false,
"status": 200,
"name": "result.pdf",
"credits": 28,
"duration": 245,
"remainingCredits": 98194839
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/rotate' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf",
"name": "result.pdf",
"angle": 90,
"pages": "0-2,4"
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf";
// Angle in degrees. Supported values are 90, 180, 270.
const Angle = 90;
// Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty
const Pages = "0-2,4";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Rotate PDF` API endpoint
var queryPath = `/v1/pdf/edit/rotate`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
url: SourceFileUrl, name: path.basename(DestinationFile), angle: Angle, pages: Pages, async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
function checkIfJobIsCompleted(jobId, resultFileUrl) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl); }, 3000);
}
else if (data.status == "success") {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf"
# Angle in degrees. Supported values are 90, 180, 270.
Angle = 90
# Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty
Pages = "0-2,4"
# Destination PDF file name
DestinationFile = ".\\result.pdf"
# (!) Make asynchronous job
Async = True
def main(args = None):
rotatePDF(SourceFileURL, DestinationFile)
def rotatePDF(uploadedFileUrl, destinationFile):
"""Rotate PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["url"] = uploadedFileUrl
parameters["name"] = os.path.basename(destinationFile)
parameters["angle"] = Angle
parameters["pages"] = Pages
parameters["async"] = Async
# Prepare URL for 'Rotate PDF' API request
url = "{}/pdf/edit/rotate".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# URL of the result file
resultFileUrl = json["url"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.IO;
using System.Net;
using Newtonsoft.Json.Linq;
using System.Threading;
using System.Collections.Generic;
using Newtonsoft.Json;
// Cloud API asynchronous "Rotate PDF" job example.
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf";
// Angle in degrees. Supported values are 90, 180, 270.
const int Angle = 90;
// Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty
const string Pages = "0-2,4";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
// (!) Make asynchronous job
const bool Async = true;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Rotate PDF` API call
string url = "https://api.pdf.co/v1/pdf/edit/rotate";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("url", SourceFileUrl);
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("angle", Angle);
parameters.Add("pages", Pages);
parameters.Add("async", Async);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Asynchronous job ID
string jobId = json["jobId"].ToString();
// URL of generated PDF file that will available after the job completion
string resultFileUrl = json["url"].ToString();
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do
{
string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status);
if (status == "success")
{
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
break;
}
else if (status == "working")
{
// Pause for a few seconds
Thread.Sleep(3000);
}
else
{
Console.WriteLine(status);
break;
}
}
while (true);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
static string CheckJobStatus(string jobId)
{
using (WebClient webClient = new WebClient())
{
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
string response = webClient.DownloadString(url);
JObject json = JObject.Parse(response);
return Convert.ToString(json["status"]);
}
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-optimize/sample.pdf";
// Angle in degrees. Supported values are 90, 180, 270.
final static Integer Angle = 90;
// Comma-separated list of page indices (or ranges) to process. Example: '0,3-5,7-'. For ALL pages just leave this param empty
final static String Pages = "0-2,4";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Rotate PDF` API call
String query = "https://api.pdf.co/v1/pdf/edit/rotate";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"url\": \"%s\", \"name\": \"%s\", \"angle\": \"%d\", \"pages\": \"%s\"}",
SourceFileUrl,
DestinationFile.getFileName(),
Angle,
Pages);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Optimize PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF Search and Delete Text
Source: https://developer.pdf.co/api/pdf-search-text-and-delete
Delete text from the PDF document with search strings.
**Try it live:** [PDF Search and Delete Text → API Tester](/api-tester/pdf-search-text-and-delete) — send a real request from your browser.
## `POST /v1/pdf/edit/delete-text`
When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `searchString` | string | *Yes* | - | Single string to search and delete. Provide either `searchString` or `searchStrings`. Text to search can support regular expressions if you set the `regex` param to true. |
| `searchStrings` | array\[string] | *Yes* | - | Array of strings to search and delete. Provide either `searchString` or `searchStrings`. |
| `redactions` | array | *No* | - | A list of redaction objects, each defining a rectangular area to remove or cover. Each redaction object includes the following fields: |
| `page` | integer | *Yes* | - | The zero-based index of the page where the redaction should be applied. For example, 0 = first page. |
| `x` | float | *Yes* | - | The X-coordinate of the top-left corner of the redaction box, in PDF units (points). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `y` | float | *Yes* | - | The Y-coordinate of the top-left corner of the redaction box, in PDF units (points). [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `width` | float | *Yes* | - | The width of the redaction area in points. [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `height` | float | *Yes* | - | The height of the redaction area in points. [Use PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to measure coordinates. |
| `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. |
| `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. |
| `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). |
| `MakeUnsearchable` | boolean | *No* | `false` | If `true`, the output PDF is made unsearchable by replacing all pages with images. If `false`, the document remains searchable with only the specified text removed. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. If not specified, the default configuration processes all pages. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `removeTextUnderPatch` | boolean | *No* | `true` | Controls whether to remove text under the patch or not |
| `usepatch` | boolean | *No* | `false` | Controls whether to use a patch or not |
| `patchColor` | string | *No* | `#000000` | Controls the color of the patch |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Showing Redacted Text
By default when we delete text using [post-tag-pdf-edit-delete-text](/api/pdf-search-text-and-delete) it will simply remove text leaving a space where the text was.
In the case where you need to blackout deleted text it can be acheived using following `profiles` parameters.
* Set `UsePatch` parameter to `true`.
* Set `PatchColor` parameter to color we want to use for redacting in `hex` format. For example: `'PatchColor': '#000000'`.
In case we want to only blackout text, but *not remove it* so that we can still copy it, we can do so using `RemoveTextUnderPatch` parameter and set it to `false`.
If `RemoveTextUnderPatch` is set to `false` then a user could still copy the text making the redaction less secure than you might require!
```
{
"profiles": "{'UsePatch': true, 'PatchColor': '#000000', 'RemoveTextUnderPatch': true}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"name": "pdfWithTextDeleted",
"caseSensitive": "false",
"searchString": "Invoice",
"replacementLimit": 0,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.us-west-2.amazonaws.com/ZOSEQZFNVCYLD5N5CJFVIYQKBVLR8OKD/pdfWithTextDeleted.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzECYaDKOO4WmO5C5shyOYYSKCAVsAo6VkB5HQjTBd9dMlJujQdEkPfNdPeLfq2mF54s2ESZBmIAJ5UgDUo3J9R475CCS4M3nuuo%2FSJwRy5gNiJdb1ZY0uCtP87x83nH%2B%2BSDu5JK%2F%2BEOrd3MREt8KE3BsQOrv%2FKMdnK%2BT5nJ2x2hC87vHue%2FudY7%2FWX54vx4tfFobEyhEozLbPnwYyKOdEsYYWH7e8tm7XV4UeKxCoKMaXSEPvOod80hR62qXnEI42fOsON3M%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHLUVIAIPX/20230220/us-west-2/s3/aws4_request&X-Amz-Date=20230220T205521Z&X-Amz-SignedHeaders=host&X-Amz-Signature=9f79c1a30d4f373e495e735e908375dad2ae6dcafcee761a477748c2b8298605",
"pageCount": 1,
"error": false,
"status": 200,
"name": "pdfWithTextDeleted.pdf",
"credits": 21,
"duration": 189,
"remainingCredits": 96235635
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/delete-text' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"name": "pdfWithTextDeleted",
"caseSensitive": "false",
"searchString": "Invoice",
"replacementLimit": 0,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Delete Text from PDF` API endpoint
var queryPath = `/v1/pdf/edit/delete-text`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, url: SourceFileUrl, searchString: 'conspicuous'
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf"
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
deleteTextFromPdf(SourceFileURL, DestinationFile)
def deleteTextFromPdf(uploadedFileUrl, destinationFile):
"""Delete Text from PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["url"] = uploadedFileUrl
parameters["searchString"] = "conspicuous"
# Prepare URL for 'Delete Text from PDF' API request
url = "{}/pdf/edit/delete-text".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("url", SourceFileUrl);
parameters.Add("searchString", "conspicuous");
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// URL of `Delete Text from PDF` API call
string url = "https://api.pdf.co/v1/pdf/edit/delete-text";
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated PDF file
string resultFileUrl = json["url"].ToString();
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "**********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Delete Text from PDF` API call
String query = "https://api.pdf.co/v1/pdf/edit/delete-text";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"conspicuous\"}",
DestinationFile.getFileName(),
Password,
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Delete Text from PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Search and Replace with Image
Source: https://developer.pdf.co/api/pdf-search-text-and-replace/image
Modify a PDF file by searching for specific text and replacing it with an image.
**Try it live:** [Search and Replace with Image → API Tester](/api-tester/pdf-search-text-and-replace/image) — send a real request from your browser.
## `POST /v1/pdf/edit/replace-text-with-image`
When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. |
| `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. |
| `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `searchString` | string | *Yes* | - | Single text replacement. Word or phrase to be replaced. Text to search can support regular expressions if you set the `regex` param to true. |
| `replaceImage` | string | *Yes* | - | Image URL or datauri to be inserted in the document |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `AutoCropImages` | boolean | *No* | - | Controls whether to crop empty space around an inserted image. See [Crop Empty Space Around Images](#crop-empty-space-around-images) for more information. |
### Crop Empty Space Around Images
If you require to crop empty space around an inserted image use the following:
```json theme={null}
{
"profiles": "{'AutoCropImages': true}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"searchString": "Your Company Name",
"caseSensitive": false,
"replaceImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png",
"pages": "0",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/7ea2b532988742508906cff59be0180e/sample.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.pdf",
"remainingCredits": 99150679,
"credits": 77
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/replace-text-with-image' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"searchString": "Your Company Name",
"caseSensitive": false,
"replaceImage": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-edit/logo.png",
"pages": "0",
"async": false
}'
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-2.pdf"
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination PDF file name
DestinationFile = ".\\result.pdf"
# (!) Make asynchronous job
Async = True
def main(args = None):
replaceImageFromPdf(SourceFileURL, DestinationFile)
def replaceImageFromPdf(uploadedFileUrl, destinationFile):
"""Replace Text With Image from PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co
parameters = {}
parameters["async"] = Async
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["url"] = uploadedFileUrl
parameters["searchString"] = "[CLIENT-SIGNATURE]"
parameters["caseSensitive"] = True
parameters["replaceImage"] = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/john-doe-signature-image.png"
# Prepare URL for 'Replace Text With Image from PDF' API request
url = "{}/pdf/edit/replace-text-with-image".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# URL of the result file
resultFileUrl = json["url"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.IO;
using System.Net;
using Newtonsoft.Json.Linq;
using System.Threading;
using System.Collections.Generic;
using Newtonsoft.Json;
// Cloud API asynchronous "Replace Text With Image from PDF" job example.
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "*****************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-2.pdf";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination PDF file name
const string DestinationFile = @".\result.pdf";
// (!) Make asynchronous job
const bool Async = true;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `Replace Text With Image from PDF` API call
string url = "https://api.pdf.co/v1/pdf/edit/replace-text-with-image";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("url", SourceFileUrl);
parameters.Add("async", Async);
parameters.Add("searchString", "[CLIENT-SIGNATURE]");
parameters.Add("caseSensitive", true);
parameters.Add("replaceImage", "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/john-doe-signature-image.png");
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Asynchronous job ID
string jobId = json["jobId"].ToString();
// URL of generated PDF file that will available after the job completion
string resultFileUrl = json["url"].ToString();
// Check the job status in a loop.
// If you don't want to pause the main thread you can rework the code
// to use a separate thread for the status checking and completion.
do
{
string status = CheckJobStatus(jobId); // Possible statuses: "working", "failed", "aborted", "success".
// Display timestamp and status (for demo purposes)
Console.WriteLine(DateTime.Now.ToLongTimeString() + ": " + status);
if (status == "success")
{
// Download PDF file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated PDF file saved as \"{0}\" file.", DestinationFile);
break;
}
else if (status == "working")
{
// Pause for a few seconds
Thread.Sleep(3000);
}
else
{
Console.WriteLine(status);
break;
}
}
while (true);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
static string CheckJobStatus(string jobId)
{
using (WebClient webClient = new WebClient())
{
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
string url = "https://api.pdf.co/v1/job/check?jobid=" + jobId;
string response = webClient.DownloadString(url);
JObject json = JObject.Parse(response);
return Convert.ToString(json["status"]);
}
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Replace Text With Image from PDF` API call
String query = "https://api.pdf.co/v1/pdf/edit/replace-text-with-image";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"/creativecommons.org/licenses/by-sa/3.0/\", \"replaceImage\": \"https://pdfco-test-files.s3.us-west-2.amazonaws.com/image-to-pdf/image1.png\"}",
DestinationFile.getFileName(),
Password,
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Replace Text With Image from PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Search and Replace with Text
Source: https://developer.pdf.co/api/pdf-search-text-and-replace/text
Replaces text in a PDF file with a new text.
**Try it live:** [Search and Replace with Text → API Tester](/api-tester/pdf-search-text-and-replace/text) — send a real request from your browser.
## `POST /v1/pdf/edit/replace-text`
When using regular expressions in JSON payloads, ensure that backslashes are properly escaped. For example, a single backslash `\` should be written as `\\`.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------------- | -------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `searchString` | string | *Yes* | - | Single string to search. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. Text to search can support regular expressions if you set the `regex` param to true. |
| `replaceString` | string | *Yes* | - | Single replacement string. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. |
| `searchStrings` | array\[string] | *Yes* | - | Array of strings to search. Provide either `searchString`/`replaceString` or `searchStrings`/`replaceStrings`. |
| `replaceStrings` | array\[string] | *Yes* | - | Array of replacement strings. Must match the order of `searchStrings`. |
| `replacementLimit` | integer | *No* | `0` | Limit the number of searches & replacements for every item. The value 0 means every found occurrence will be replaced. |
| `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. |
| `regex` | boolean | *No* | `false` | Set to `true` to use regular expression for search string(s). |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `AutoCropImages` | boolean | *No* | `false` | If you require to crop empty space around an inserted image use the following: `profiles": { 'AutoCropImages': true }` |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `YAdjustmentForReplacementText` | integer | *No* | - | Adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. See [Adjust Text Alignment](#adjust-text-alignment) for more details. |
### Adjust Text Alignment
Users may have encountered an issue when using this API endpoint to replace text in a **PDF** document. The replaced text might appear slightly higher than the original text or the surrounding text, causing alignment issues.
To fix this issue, we have added a new parameter called `YAdjustmentForReplacementText` in the `profiles` parameter of the API request. This parameter allows you to adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. Negative values for this parameter move text up, positive values move text down.
Here’s an example of how to use the `YAdjustmentForReplacementText` parameter. In this example API request, the `YAdjustmentForReplacementText` parameter has been set to `-1`, which moves the replaced text `1` unit up vertically, resulting in better alignment with the original text.
```json theme={null}
{
"profiles": "{'YAdjustmentForReplacementText': '-1'}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-search-and-replace/sample-agreement-template-signature-page-1.pdf",
"searchStrings": [
"[CLIENT-NAME]",
"[CLIENT-COMPANY]"
],
"replaceStrings": [
"John Doe",
"Skynet 3000"
],
"caseSensitive": true,
"replacementLimit": 1,
"pages": "",
"password": "",
"name": "finalFile",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/e79f0b9c82984740973ca670d7c93cad/finalFile.pdf",
"pageCount": 1,
"error": false,
"status": 200,
"name": "finalFile.pdf",
"remainingCredits": 99089875,
"credits": 21
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/edit/replace-text' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"name": "pdfWithTextReplaced",
"caseSensitive": "false",
"searchString": "Your Company Name",
"replaceString": "Acme ltd.",
"replacementLimit": 0,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination PDF file name
const DestinationFile = "./result.pdf";
// Prepare request to `Replace Text from PDF` API endpoint
var queryPath = `/v1/pdf/edit/replace-text`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), password: Password, url: SourceFileUrl, searchString: 'Your Company Name', replaceString: 'XYZ LLC'
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download PDF file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Direct URL of source PDF file.
# You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
SourceFileURL = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf"
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination PDF file name
DestinationFile = ".\\result.pdf"
def main(args = None):
replaceStringFromPdf(SourceFileURL, DestinationFile)
def replaceStringFromPdf(uploadedFileUrl, destinationFile):
"""Replace Text from PDF using PDF.co Web API"""
# Prepare requests params as JSON
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["url"] = uploadedFileUrl
parameters["searchString"] = "Your Company Name"
parameters["replaceString"] = "XYZ LLC"
# Prepare URL for 'Replace Text from PDF' API request
url = "{}/pdf/edit/replace-text".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
if __name__ == '__main__':
main()
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination PDF file name
final static Path DestinationFile = Paths.get(".\\result.pdf");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `Replace Text from PDF` API call
String query = "https://api.pdf.co/v1/pdf/edit/replace-text";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"url\": \"%s\", \"searchString\": \"Your Company Name\", \"replaceString\": \"XYZ LLC\"}",
DestinationFile.getFileName(),
Password,
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated PDF file
String resultFileUrl = json.get("url").getAsString();
// Download PDF file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated PDF file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
Cloud API asynchronous "Replace Text from PDF" job example (allows to avoid timeout errors).
" . date(DATE_RFC2822) . ": " . $status . "";
if ($status == "success")
{
// Display link to the file with conversion results
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# Split PDF
Source: https://developer.pdf.co/api/pdf-split/by-pages
Split a PDF into multiple files by specifying page numbers or page ranges to keep in each output.
**Try it live:** [Split PDF → API Tester](/api-tester/pdf-split/by-pages) — send a real request from your browser.
## `POST /v1/pdf/split`
When splitting a document the pages parameter controls which `pages` to split out into individual documents. The page limit should not exceed the number of pages in the document - for example, you cannot split a 100 page document into 200 individual documents, however you can split it into 100 individual documents.The `pages` parameter is 1-based, meaning the first page is `1` and not `0`.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify pages as comma-separated page numbers and ranges to process (e.g. "1, 2, 5-10" or "3-" for page 3 to the end). The first-page index is 1. Use "!" before a number for inverted page numbers (e.g. "!1" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `fixed_output_filename` | boolean | *No* | `false` | Defines whether the page range is appended to the output filename. When set to `true`, the output filename remains exactly as specified in the `name` parameter. When set to `false` (default), the filename automatically includes the page range of the extracted pages for each output file. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf",
"pages": "1-2,3-",
"inline": true,
"name": "result.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/1e9a7f2c46834160903276716424382b/result_page1-2.pdf",
"https://pdf-temp-files.s3.amazonaws.com/c976b9f89a2e460786a3d5c0deeeef67/result_page3-4.pdf"
],
"pageCount": 4,
"error": false,
"status": 200,
"name": "result.pdf",
"remainingCredits": 98441
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/split' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/sample.pdf",
"pages": "1-2,3-",
"inline": true,
"name": "result.pdf",
"async": false
}'
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'.
Pages = "1-2,3-"
# (!) Make asynchronous job
Async = True
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
splitPDF(uploadedFileUrl)
def splitPDF(uploadedFileUrl):
"""Split PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["async"] = Async
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'Split PDF' API request
url = "{}/pdf/split".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# URL of the result file
resultFilePlaceholder = json["url"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
resJsonImgFiles = requests.get(resultFilePlaceholder)
# Download generated PNG files
part = 1
for resultFileUrl in resJsonImgFiles.json():
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.pdf"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file to split
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'.
const string Pages = "1-2,3-";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. SPLIT UPLOADED PDF
// URL for `Split PDF` API call
var url = "https://api.pdf.co/v1/pdf/split";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated PDF files
int part = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\part{0}.pdf", part);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
part++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file to split
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page numbers (or ranges) to process. Example: '1,3-5,7-'.
final static String Pages = "1-2,3-";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. SPLIT UPLOADED PDF
SplitPdf(webClient, API_KEY, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void SplitPdf(OkHttpClient webClient, String apiKey, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `Split PDF` API call
String query = "https://api.pdf.co/v1/pdf/split";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"pages\": \"%s\", \"url\": \"%s\"}",
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated PDF files
JsonArray urls = json.get("urls").getAsJsonArray();
int part = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\part%s.pdf", part);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Splitted part saved as \"%s\".", localFileName));
part++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF Splitting Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display request error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# Split PDF by Text Search
Source: https://developer.pdf.co/api/pdf-split/by-text-search-or-barcode
Split a PDF into multiple files at every page that matches a text search or barcode pattern.
**Try it live:** [Split PDF by Text Search → API Tester](/api-tester/pdf-split/by-text-search-or-barcode) — send a real request from your browser.
## `POST /v1/pdf/split2`
This endpoint decides where to split using `searchString` only. Split points are every page that matches the search, so there is no page-selection parameter to configure.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `searchString` | string | *Yes* | - | Text to search for on pages. Must be a string. |
| `regexSearch` | boolean | *No* | `false` | Set to true to enable regular expression search for the `searchString(s)` parameter. |
| `caseSensitive` | boolean | *No* | `true` | Set to `false` to don't use case-sensitive search. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `excludeKeyPages` | boolean | *No* | `false` | Set to true to exclude pages where the searchString text was found. |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
### searchString
Text to search for on pages. Must be a string.
To search for a barcode use the following macros string: `[[barcode:]]`.
To search for barcode type without analyzing its value, use this notation instead: `[[barcode:]].`
Example #1, split by QR code: "searchString": "\[\[barcode:qrcode]]".
Example #2, split by QR code with value: "searchString": "\[\[barcode:qrcode pdfco]]".
Example #3, split by QR code with value search with regex: "searchString": "\[\[barcode:qrcode /pdf.co/]]".
Example #4, split by QR code or datamatrix with value search with regex: "searchString": "\[\[barcode:qrcode,datamatrix /pdf.co/]]".
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/split_by_barcode.pdf",
"searchString": "[[barcode:qrcode,datamatrix /pdf\\.co/]]",
"excludeKeyPages": true,
"regexSearch": false,
"caseSensitive": false,
"inline": true,
"name": "output-split-by-barcode",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.us-west-2.amazonaws.com/A2WX2GR0PX4818EIKW96VR3BZTK5FWT2/output-split-by-barcode_page1.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=8241ad05ecb5555cbbd4998b5c334104f2849bf4177384e86fbb5cc5d7e81ce8",
"https://pdf-temp-files.s3.us-west-2.amazonaws.com/B6Z9J274GZ5BK5QYK547ST4T5WF61LNQ/output-split-by-barcode_page3-5.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=94764cfb37819f2a4885ba064dd1ae20f38f42d6bc6c1a208010637fca74a591",
"https://pdf-temp-files.s3.us-west-2.amazonaws.com/XT5TD1BDBFDNKX0LM6N5GLFLOAF1UC0Y/output-split-by-barcode_page7-9.pdf?X-Amz-Expires=3600&X-Amz-Security-Token=FwoGZXIvYXdzEK3%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaDH1Gv1Q88EtgGpfAYiKCAaQTLV5ot8KMblEXIEFzeznT8mOeGKylp0uktJk2Se8SK5r3nfQTJKa8JqJE0GcW9vOtcBPPqHcPZXf2iQkvSk3yvFJv6cDj8%2B6kck0Eadz4BOXz0ljrE1Vt%2BX2gItx86Fd8rldFG3TL7u99FKiuc1rN9OaBRJpPHL12fVP2gjuVUUIomqShmQYyKHbhGDuLKoCWq%2BdLkggz2eTJna6w9eWR7QMvpIJxc8sBGFT1WEm%2FsyA%3D&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIA4NRRSZPHORHIVCFW/20220919/us-west-2/s3/aws4_request&X-Amz-Date=20220919T114402Z&X-Amz-SignedHeaders=host&X-Amz-Signature=0a7c90a05fd159659451d29273284fbf422d34bd204c07fbc9abdf7a36a84294"
],
"pageCount": 10,
"error": false,
"status": 200,
"name": "output-split-by-barcode.pdf",
"credits": 350,
"duration": 4456,
"remainingCredits": 98221710
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/split2' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-split/split_by_barcode.pdf",
"searchString": "[[barcode:qrcode,datamatrix /pdf\\.co/]]",
"excludeKeyPages": true,
"regexSearch": false,
"caseSensitive": false,
"inline": true,
"name": "output-split-by-barcode",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file to split
const SourceFile = "./sample.pdf";
// Split Search String
const SplitText = "invoice number";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. SPLIT UPLOADED PDF
splitPdf(API_KEY, uploadedFileUrl, SplitText);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + err);
}
});
});
});
}
function splitPdf(apiKey, uploadedFileUrl, splitText) {
// Prepare request to `Split PDF By Text` API endpoint
var queryPath = `/v1/pdf/split2`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
searchString: splitText, url: uploadedFileUrl, async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
console.log(`Job #${data.jobId} has been created!`);
checkIfJobIsCompleted(data.jobId, data.url);
}
else {
// Service reported error
console.log("splitPdf(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("splitPdf(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
function checkIfJobIsCompleted(jobId, resultFileUrlJson) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrlJson) }, 3000);
}
else if (data.status == "success") {
request({ method: 'GET', uri: resultFileUrlJson, gzip: true },
function (error, response, body) {
// Parse JSON response
let respJsonFileArray = JSON.parse(body);
let part = 1;
respJsonFileArray.forEach((url) => {
var localFileName = `./part${part}.pdf`;
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated PDF file saved as "${localFileName} file."`);
});
});
part++;
}, this);
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
import time
import datetime
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Split by Text
SplitText = "invoice number"
# (!) Make asynchronous job
Async = True
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
splitPDF(uploadedFileUrl)
def splitPDF(uploadedFileUrl):
"""Split PDF using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["async"] = Async
parameters["searchString"] = SplitText
parameters["url"] = uploadedFileUrl
# Prepare URL for 'Split PDF By Text' API request
url = "{}/pdf/split2".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Asynchronous job ID
jobId = json["jobId"]
# URL of the result file
resultFilePlaceholder = json["url"]
# Check the job status in a loop.
# If you don't want to pause the main thread you can rework the code
# to use a separate thread for the status checking and completion.
while True:
status = checkJobStatus(jobId) # Possible statuses: "working", "failed", "aborted", "success".
# Display timestamp and status (for demo purposes)
print(datetime.datetime.now().strftime("%H:%M.%S") + ": " + status)
if status == "success":
resJsonImgFiles = requests.get(resultFilePlaceholder)
# Download generated files
part = 1
for resultFileUrl in resJsonImgFiles.json():
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.pdf"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
break
elif status == "working":
# Pause for a few seconds
time.sleep(3)
else:
print(status)
break
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def checkJobStatus(jobId):
"""Checks server job status"""
url = f"{BASE_URL}/job/check?jobid={jobId}"
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
return json["status"]
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file to split
const string SourceFile = @".\sample.pdf";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. SPLIT UPLOADED PDF By Text
// URL for `Split PDF By Text` API call
var url = "https://api.pdf.co/v1/pdf/split2";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("searchString", "invoice number");
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated PDF files
int part = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\part{0}.pdf", part);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
part++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file to split
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Split By Text
final static String SplitText = "invoice number";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. SPLIT UPLOADED PDF
SplitPdf(webClient, API_KEY, SplitText, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void SplitPdf(OkHttpClient webClient, String apiKey, String splitText, String uploadedFileUrl) throws IOException
{
// Prepare URL for `Split PDF By Text` API call
String query = "https://api.pdf.co/v1/pdf/split2";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"searchString\": \"%s\", \"url\": \"%s\"}",
splitText,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated PDF files
JsonArray urls = json.get("urls").getAsJsonArray();
int part = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\part%s.pdf", part);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Splitted part saved as \"%s\".", localFileName));
part++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF Splitting Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display request error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to CSV
Source: https://developer.pdf.co/api/pdf-to-csv
Convert PDF and scanned images into CSV representation with layout, columns, rows, and tables.
**Try it live:** [PDF to CSV → API Tester](/api-tester/pdf-to-csv) — send a real request from your browser.
## `POST /v1/pdf/convert/to/csv`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. |
| `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | `None` | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file.
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `body` | string | Stringified CSV content |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf",
"lang": "eng",
"inline": "true",
"unwrap": "",
"pages": "0-",
"rect": "",
"async": "false",
"name": "result.csv",
"password": "",
"lineGrouping": "",
"profiles": ""
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": "\"Your Company Name\",\"\",\"\",\"\",\r\n\"Your Address\",\"\",\"\",\"\",\r\n\"City, State Zip\",\"\",\"\",\"\",\r\n\"\",\"\",\"\",\"Invoice No. 123456\",\r\n\"\",\"\",\"\",\"Invoice Date 01/01/2016\",\r\n\"Client Name\",\"\",\"\",\"\",\r\n\"Address\",\"\",\"\",\"\",\r\n\"City, State Zip\",\"\",\"\",\"\",\r\n\"Notes\",\"\",\"\",\"\",\r\n\"Item\",\"Quantity\",\"Price\",\"Total\",\r\n\"Item 1\",\"1\",\"40.00\",\"40.00\",\r\n\"Item 2\",\"2\",\"30.00\",\"60.00\",\r\n\"Item 3\",\"3\",\"20.00\",\"60.00\",\r\n\"Item 4\",\"4\",\"10.00\",\"40.00\",\r\n\"\",\"\",\"TOTAL\",\"200.00\",\r\n",
"pageCount": 2,
"error": false,
"status": 200,
"name": "result.csv",
"remainingCredits": 616411,
"credits": 56
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/csv' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-csv/sample.pdf",
"lang": "eng",
"inline": "true",
"unwrap": "",
"pages": "0-",
"rect": "",
"async": "false",
"name": "result.csv",
"password": "",
"lineGrouping": "",
"profiles": ""
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "*********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination CSV file name
const DestinationFile = "./result.csv";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO CSV
convertPdfToCsv(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToCsv(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To CSV` API endpoint
var queryPath = `/v1/pdf/convert/to/csv`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl, async: true
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Job #${data.jobId} has been created!`);
if (data.error == false) {
checkIfJobIsCompleted(data.jobId, data.url, destinationFile);
}
else {
// Service reported error
console.log("convertPdfToCsv(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToCsv(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
function checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile) {
let queryPath = `/v1/job/check`;
// JSON payload for api request
let jsonPayload = JSON.stringify({
jobid: jobId
});
let reqOptions = {
host: "api.pdf.co",
path: queryPath,
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
console.log(`Checking Job #${jobId}, Status: ${data.status}, Time: ${new Date().toLocaleString()}`);
if (data.status == "working") {
// Check again after 3 seconds
setTimeout(function () { checkIfJobIsCompleted(jobId, resultFileUrl, destinationFile); }, 3000);
}
else if (data.status == "success") {
// Download CSV file
var file = fs.createWriteStream(destinationFile);
https.get(resultFileUrl, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated CSV file saved as "${destinationFile}" file.`);
});
});
}
else {
console.log(`Operation ended with status: "${data.status}".`);
}
})
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination CSV file name
DestinationFile = ".\\result.csv"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToCSV(uploadedFileUrl, DestinationFile)
def convertPdfToCSV(uploadedFileUrl, destinationFile):
"""Converts PDF To CSV using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To CSV' API request
url = "{}/pdf/convert/to/csv".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "**************************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination CSV file name
const string DestinationFile = @".\result.csv";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["status"].ToString() != "error")
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO CSV
// URL for `PDF To CSV` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/csv";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["status"].ToString() != "error")
{
// Get URL of generated CSV file
string resultFileUrl = json["url"].ToString();
// Download CSV file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated CSV file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination CSV file name
final static Path DestinationFile = Paths.get(".\\result.csv");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
String status = json.get("status").getAsString();
if (!status.equals("error"))
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO CSV
PdfToCsv(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToCsv(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To CSV` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/csv";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
String status = json.get("status").getAsString();
if (!status.equals("error"))
{
// Get URL of generated CSV file
String resultFileUrl = json.get("url").getAsString();
// Download CSV file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated CSV file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To CSV Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to XLS
Source: https://developer.pdf.co/api/pdf-to-excel/xls
Convert PDF to Excel(.xls) with layout and fonts preserved.
**Try it live:** [PDF to XLS → API Tester](/api-tester/pdf-to-excel/xls) — send a real request from your browser.
## `POST /v1/pdf/convert/to/xls`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file.
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/60c6b9f50280495a9567f73a0a394252/sample.xlsx",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.xlsx",
"remainingCredits": 60568
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination XLS file name
const DestinationFile = "./result.xls";
// Prepare request to `PDF To XLS` API endpoint
var queryPath = `/v1/pdf/convert/to/xls`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download XLS file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated XLS file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const string SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination XLS file name
const string DestinationFile = @".\result.xls";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// URL for `PDF To XLS` API call
string url = "https://api.pdf.co/v1/pdf/convert/to/xls";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", SourceFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
try
{
// Execute POST request with JSON payload
string response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated XLS file
string resultFileUrl = json["url"].ToString();
// Download XLS file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated XLS file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
final static String SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination XLS file name
final static Path DestinationFile = Paths.get(".\\result.xls");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// Prepare URL for `PDF To XLS` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/xls";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
DestinationFile.getFileName(),
Password,
Pages,
SourceFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated XLS file
String resultFileUrl = json.get("url").getAsString();
// Download XLS file
downloadFile(webClient, resultFileUrl, DestinationFile.toFile());
System.out.printf("Generated XLS file saved as \"%s\" file.", DestinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
PDF To Excel Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF to XLSX
Source: https://developer.pdf.co/api/pdf-to-excel/xlsx
Convert PDF to Excel(.xlsx) with layout and fonts preserved.
**Try it live:** [PDF to XLSX → API Tester](/api-tester/pdf-to-excel/xlsx) — send a real request from your browser.
## `POST /v1/pdf/convert/to/xlsx`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `ColumnDetectionMode` | string | *No* | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file.
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/60c6b9f50280495a9567f73a0a394252/sample.xlsx",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.xlsx",
"remainingCredits": 60568
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/xlsx?=' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Direct URL of source PDF file.
// You can also upload your own file into PDF.co and use it as url. Check "Upload File" samples for code snippets: https://github.com/pdfdotco/pdf-co-api-samples/tree/master/File%20Upload/
const SourceFileUrl = "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-excel/sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination XLSX file name
const DestinationFile = "./result.xlsx";
// Prepare request to `PDF To XLSX` API endpoint
var queryPath = `/v1/pdf/convert/to/xlsx`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(DestinationFile), password: Password, pages: Pages, url: SourceFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
// Parse JSON response
var data = JSON.parse(d);
if (data.error == false) {
// Download XLSX file
var file = fs.createWriteStream(DestinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated XLSX file saved as "${DestinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log(data.message);
}
});
}).on("error", (e) => {
// Request error
console.log(e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination Excel file name
DestinationFile = ".\\result.xlsx"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToExcel(uploadedFileUrl, DestinationFile)
def convertPdfToExcel(uploadedFileUrl, destinationFile):
"""Converts PDF To Excel using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Xlsx' API request
url = "{}/pdf/convert/to/xlsx".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination XLSX file name
const string DestinationFile = @".\result.xlsx";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO XLSX
// URL for `PDF To XLSX` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/xlsx";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated XLSX file
string resultFileUrl = json["url"].ToString();
// Download XLSX file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated XLSX file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination XLSX file name
final static Path DestinationFile = Paths.get(".\\result.xlsx");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO XLSX
PdfToXlsx(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToXlsx(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To XLSX` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/xlsx";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated XLSX file
String resultFileUrl = json.get("url").getAsString();
// Download XLSX file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated XLSX file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Excel Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
return $status;
}
?>
```
# PDF to HTML
Source: https://developer.pdf.co/api/pdf-to-html
Convert PDF and scanned images into HTML representation with text, fonts, images, vectors, formatting preserved.
**Try it live:** [PDF to HTML → API Tester](/api-tester/pdf-to-html) — send a real request from your browser.
## `POST /v1/pdf/convert/to/html`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `OptimizeImages` | boolean | *No* | `true` | Some PDF may have high quality images used in the document and you may need to keep the quality of these images in the output HTML. By default PDF to HTML is optimizing images and you can easily turn it off. See [Control Image Quality](#control-image-quality) for more information. |
| `OutputPageWidth` | integer | *No* | `1024` | Control page width (in pixels) for output HTML. Height is calculated and used according to the original pdf pages ratio. See [Control Output Page Width](#control-output-page-width) for more information. |
| `AdditionalCssStyles` | string | *No* | \`\` | To inject CSS for layout options in your HTML. Example: `#canvas { zoom: 50%; }`. Scale the div that contains all generated HTML pages by 50%. See [Inject CSS](#inject-css) for more information. |
| `saveImages` | integer | *No* | - | Controls whether to save images in the output HTML. See [Disable Images](#disable-images) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Disable Images
To turn off images output set the following profile:
```json theme={null}
{
"profiles": "{ 'saveImages': 0 }"
}
```
### Control Image Quality
Some **PDF** may have high quality images used in the document and you may need to keep the quality of these images in the output **HTML**. By default [PDF to HTML](/api/pdf-to-html) is optimizing images and you can easily turn it off with the following profile:
```json theme={null}
{
"profiles": "{ 'OptimizeImages': false }"
}
```
### Control Output Page Width
Control page width output as follows:
```json theme={null}
{
"profiles": "{ 'OutputPageWidth': 2048 }"
}
```
### Inject CSS
To inject CSS for layout options in your HTML use the following:
```json theme={null}
{
"profiles": "{ 'AdditionalCssStyles': '#canvas { zoom: 50%; }' }"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf",
"inline": false,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://pdf-temp-files.s3.amazonaws.com/a7a86e9f29f84f5180624bdec1facfc2/index.html",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.html",
"remainingCredits": 60110
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/html' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-html/sample.pdf",
"inline": false,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination HTML file name
const DestinationFile = "./result.html";
// Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design.
const PlainHtml = false;
// Set to `true` if your document has the column layout like a newspaper.
const ColumnLayout = false;
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO HTML
convertPdfToHtml(API_KEY, uploadedFileUrl, Password, Pages, PlainHtml, ColumnLayout, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToHtml(apiKey, uploadedFileUrl, password, pages, plainHtml, columnLayout, destinationFile) {
// Prepare request to `PDF To HTML` API endpoint
var queryPath = `/v1/pdf/convert/to/html`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, simple: plainHtml, columns: columnLayout, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download HTML file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated HTML file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToHtml(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToHtml(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "***************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination Html file name
DestinationFile = ".\\result.html"
# Set to $true to get simplified HTML without CSS. Default is the rich HTML keeping the document design.
PlainHtml = False
# Set to $true if your document has the column layout like a newspaper.
ColumnLayout = False
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToHtml(uploadedFileUrl, DestinationFile)
def convertPdfToHtml(uploadedFileUrl, destinationFile):
"""Converts PDF To Html using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-html
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["simple"] = PlainHtml
parameters["columns"] = ColumnLayout
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Html' API request
url = "{}/pdf/convert/to/html".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination HTML file name
const string DestinationFile = @".\result.html";
// Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design.
const bool PlainHtml = false;
// Set to `true` if your document has the column layout like a newspaper.
const bool ColumnLayout = false;
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have the direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO HTML
// URL for `PDF To HTML` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/html";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("simple", PlainHtml);
parameters.Add("columns", ColumnLayout);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated HTML file
string resultFileUrl = json["url"].ToString();
// Download HTML file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated HTML file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination HTML file name
final static Path DestinationFile = Paths.get(".\\result.html");
// Set to `true` to get simplified HTML without CSS. Default is the rich HTML keeping the document design.
final static boolean PlainHtml = false;
// Set to `true` if your document has the column layout like a newspaper.
final static boolean ColumnLayout = false;
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO HTML
PdfToHtml(webClient, API_KEY, DestinationFile, Password, Pages, PlainHtml, ColumnLayout, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToHtml(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, boolean plainHtml, boolean columnLayout, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To HTML` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/html";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"simple\": \"%s\", \"columns\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
plainHtml,
columnLayout,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated HTML file
String resultFileUrl = json.get("url").getAsString();
// Download HTML file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated HTML file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To HTML Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to JPG
Source: https://developer.pdf.co/api/pdf-to-image/jpg
PDF to high quality JPEG image conversion. High quality rendering. Also works great for thumbnail generation and previews.
**Try it live:** [PDF to JPG → API Tester](/api-tester/pdf-to-image/jpg) — send a real request from your browser.
## `POST /v1/pdf/convert/to/jpg`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. |
| `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not |
| `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not |
| `RenderCurveVectorObjects` | boolean | *No* | `true` | Render curve vector objects or not |
| `TextSmoothingMode` | string | *No* | - | Controls text smoothing mode. Available options: `HighSpeed`, `HighQuality`. |
| `VectorSmoothingMode` | string | *No* | - | Controls vector smoothing mode. Available options: `HighSpeed`, `HighQuality`. |
| `ImageInterpolationMode` | string | *No* | - | Controls image interpolation mode. Available options: `HighSpeed`, `HighQuality`. |
| `JPEGQuality` | integer | *No* | 80 | Range from `0` (lowest) to `100` (highest), default is `80`. See [profiles.JPEGQuality](#profiles-jpegquality) |
| `TIFFCompression` | string | *No* | - | Controls TIFF compression. Available options: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE`. |
| `RotateFlipType` | string | *No* | - | Controls rotation and flip type. Available options: `RotateNoneFlipNone`, `Rotate90FlipNone`, `Rotate180FlipNone`, `Rotate270FlipNone`, `RotateNoneFlipX`, `Rotate90FlipX`, `Rotate180FlipX`, `Rotate270FlipX`, `RotateNoneFlipY`, `Rotate90FlipY`, `Rotate180FlipY`, `Rotate270FlipY`, `RotateNoneFlipXY`, `Rotate90FlipXY`, `Rotate180FlipXY`, `Rotate270FlipXY`. |
| `ImageBitsPerPixel` | string | *No* | - | Controls image bits per pixel. Available options: `BPP1`, `BPP8`, `BPP24`, `BPP32`. |
| `OneBitConversionAlgorithm` | string | *No* | - | Controls one-bit conversion algorithm. Available options: `BayerOrderedDithering`, `OtsuThreshold`. |
| `FontHintingMode` | string | *No* | - | Controls font hinting mode. Available options: `Default`, `Stronger`. |
| `ResolutionOverride` | float | *No* | - | Overrides the default resolution. Specified in DPI. |
| `NightMode` | boolean | *No* | `false` | Enables night mode rendering. |
| `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Disable Text Layer
We can turn off the text layer for our render as follows:
```json theme={null}
{
"profiles": "{ 'RenderTextObjects': false }"
}
```
### Set Image Resolution
By default the screen resolution is 120 DPI. To change the rendering resolution, please use:
```json theme={null}
{
"profiles": "{ 'RenderingResolution': 300 }"
}
```
### `JPEGQuality`
To set image quality (from `0` (lowest) to `100` (highest), default is `80`) please use:
```
{
"profiles": "{ 'JPEGQuality': 75 }"
}
```
### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more.
Profiles section for more.
```json theme={null}
{
"profiles": "{
'RenderTextObjects': false,
'RenderVectorObjects': true,
'RenderImageObjects': true
}"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png",
"https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png"
],
"pageCount": 2,
"error": false,
"status": 200,
"name": "sample.png",
"duration": 1121,
"remainingCredits": 98722216,
"credits": 30
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg");
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
// `imageType` should correspond to the API we want to use,
// i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff
// we just take the last part of the path, the file extension
function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) {
// Prepare URL for PDF to Image API call
var queryPath = `/v1/pdf/convert/to/${imageType}`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download generated image files
var page = 1;
data.urls.forEach((url) => {
var localFileName = `./page${page}.${imageType}`;
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated image file saved as "${localFileName}" file.`);
});
});
page++;
}, this);
}
else {
// Service reported error
console.log("convertPDFToImage(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPDFToImage(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPDFToImage(uploadedFileUrl, "jpg")
def convertPDFToImage(uploadedFileUrl, imageType):
"""Converts PDF To Image using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-image/jpg
parameters = {}
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for PDF To Image API request
url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Download generated JPG files
part = 1
for resultFileUrl in json["urls"]:
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.{imageType}"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
const string imageType = "jpg";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have the direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
// Get URL of uploaded file to use with later API calls
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
// Prepare URL for PDF To Image API call
string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated JPEG files
int page = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\page{0}.{1}", page, imageType);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
page++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg");
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException
{
// Prepare URL for PDF To Image API call
String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType);
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated JPEG files
JsonArray urls = json.get("urls").getAsJsonArray();
int page = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\page%s.%s", page, imageType);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Downloaded \"%s\".", localFileName));
page++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Image Results
Status code: " . $status_code . "";
echo "
";
}
else
{
// JPEG and PNG formats are single-page, so the results are multiple
$resultFiles = $json["urls"];
foreach ($resultFiles as &$resultFileUrl)
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to PNG
Source: https://developer.pdf.co/api/pdf-to-image/png
PDF to high quality PNG image conversion. High quality rendering. Also works great for thumbnail generation and previews.
**Try it live:** [PDF to PNG → API Tester](/api-tester/pdf-to-image/png) — send a real request from your browser.
## `POST /v1/pdf/convert/to/png`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ------------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. |
| `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not |
| `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not |
| `RenderCurveVectorObjects` | boolean | *No* | `true` | Render curve vector objects or not |
| `TextSmoothingMode` | string | *No* | - | Controls text smoothing mode. Available options: `HighSpeed`, `HighQuality`. |
| `VectorSmoothingMode` | string | *No* | - | Controls vector smoothing mode. Available options: `HighSpeed`, `HighQuality`. |
| `ImageInterpolationMode` | string | *No* | - | Controls image interpolation mode. Available options: `HighSpeed`, `HighQuality`. |
| `TIFFCompression` | string | *No* | - | Controls TIFF compression. Available options: `None`, `LZW`, `CCITT3`, `CCITT4`, `RLE`. |
| `RotateFlipType` | string | *No* | - | Controls rotation and flip type. Available options: `RotateNoneFlipNone`, `Rotate90FlipN one`, `Rotate180FlipNone`, `Rotate270FlipNone`, `RotateNoneFlipX`, `Rotate90FlipX`, `Rotate180FlipX`, `Rotate270FlipX`, `RotateNoneFlipY`, `Rotate90FlipY`, `Rotate180FlipY`, `Rotate270FlipY`, `RotateNoneFlipXY`, `Rotate90FlipXY`, `Rotate180FlipXY`, `Rotate270FlipXY`. |
| `ImageBitsPerPixel` | string | *No* | - | Controls image bits per pixel. Available options: `BPP1`, `BPP8`, `BPP24`, `BPP32`. |
| `OneBitConversionAlgorithm` | string | *No* | - | Controls one-bit conversion algorithm. Available options: `BayerOrderedDithering`, `FloydSteinbergDithering`, `Threshold`, `OrderedDithering`. |
| `FontHintingMode` | string | *No* | - | Controls font hinting mode. Available options: `Default`, `Stronger`. |
| `ResolutionOverride` | float | *No* | - | Overrides the default resolution. Specified in DPI. |
| `NightMode` | boolean | *No* | `false` | Enables night mode rendering. |
| `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Disable Text Layer
We can turn off the text layer for our render as follows:
```json theme={null}
{
"profiles": "{ 'RenderTextObjects': false }"
}
```
### Set Image Resolution
By default the screen resolution is 120 DPI. To change the rendering resolution, please use:
```json theme={null}
{
"profiles": "{ 'RenderingResolution': 300 }"
}
```
### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more.
Profiles section for more.
```
"profiles":
{
'RenderTextObjects': false,
'RenderVectorObjects': true,
'RenderImageObjects': true
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png",
"https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png"
],
"pageCount": 2,
"error": false,
"status": 200,
"name": "sample.png",
"duration": 1121,
"remainingCredits": 98722216,
"credits": 30
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg");
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
// `imageType` should correspond to the API we want to use,
// i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff
// we just take the last part of the path, the file extension
function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) {
// Prepare URL for PDF to Image API call
var queryPath = `/v1/pdf/convert/to/${imageType}`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download generated image files
var page = 1;
data.urls.forEach((url) => {
var localFileName = `./page${page}.${imageType}`;
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated image file saved as "${localFileName}" file.`);
});
});
page++;
}, this);
}
else {
// Service reported error
console.log("convertPDFToImage(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPDFToImage(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPDFToImage(uploadedFileUrl, "jpg")
def convertPDFToImage(uploadedFileUrl, imageType):
"""Converts PDF To Image using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-image/png
parameters = {}
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for PDF To Image API request
url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Download generated JPG files
part = 1
for resultFileUrl in json["urls"]:
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.{imageType}"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
const string imageType = "jpg";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have the direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
// Get URL of uploaded file to use with later API calls
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
// Prepare URL for PDF To Image API call
string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated JPEG files
int page = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\page{0}.{1}", page, imageType);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
page++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg");
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException
{
// Prepare URL for PDF To Image API call
String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType);
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated JPEG files
JsonArray urls = json.get("urls").getAsJsonArray();
int page = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\page%s.%s", page, imageType);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Downloaded \"%s\".", localFileName));
page++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Image Results
Status code: " . $status_code . "";
echo "
";
}
else
{
// JPEG and PNG formats are single-page, so the results are multiple
$resultFiles = $json["urls"];
foreach ($resultFiles as &$resultFileUrl)
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to TIFF
Source: https://developer.pdf.co/api/pdf-to-image/tiff
PDF to high quality TIFF image conversion. High quality rendering. Also works great for thumbnail generation and previews.
**Try it live:** [PDF to TIFF → API Tester](/api-tester/pdf-to-image/tiff) — send a real request from your browser.
## `POST /v1/pdf/convert/to/tiff`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `TIFFCompression` | string | *No* | `LZW` | See [profiles.TIFFCompression](#profiles-tiffcompression) |
| `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. |
| `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Disable Text Layer
We can turn off the text layer for our render as follows:
```json theme={null}
{
"profiles": "{ 'RenderTextObjects': false }"
}
```
### Set Image Resolution
By default the screen resolution is 120 DPI. To change the rendering resolution, please use:
```json theme={null}
{
"profiles": "{ 'RenderingResolution': 300 }"
}
```
### `profiles.TIFFCompression`
TIFF has a variety of options as follows:
```
{
"profiles": "{
'RenderTextObjects': true, // Valid values: true, false
'RenderVectorObjects': true, // Valid values: true, false
'RenderImageObjects': true, // Valid values: true, false
'TIFFCompression': 'LZW', // Valid values: 'None', 'LZW', 'CCITT3', 'CCITT4', 'RLE'
}"
}
```
### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more.
Profiles section for more.
```
"profiles":
{
'RenderTextObjects': false,
'RenderVectorObjects': true,
'RenderImageObjects': true
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png",
"https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png"
],
"pageCount": 2,
"error": false,
"status": 200,
"name": "sample.png",
"duration": 1121,
"remainingCredits": 98722216,
"credits": 30
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg");
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
// `imageType` should correspond to the API we want to use,
// i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff
// we just take the last part of the path, the file extension
function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) {
// Prepare URL for PDF to Image API call
var queryPath = `/v1/pdf/convert/to/${imageType}`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download generated image files
var page = 1;
data.urls.forEach((url) => {
var localFileName = `./page${page}.${imageType}`;
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated image file saved as "${localFileName}" file.`);
});
});
page++;
}, this);
}
else {
// Service reported error
console.log("convertPDFToImage(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPDFToImage(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPDFToImage(uploadedFileUrl, "jpg")
def convertPDFToImage(uploadedFileUrl, imageType):
"""Converts PDF To Image using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-image/tiff
parameters = {}
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for PDF To Image API request
url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Download generated JPG files
part = 1
for resultFileUrl in json["urls"]:
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.{imageType}"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
const string imageType = "jpg";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have the direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
// Get URL of uploaded file to use with later API calls
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
// Prepare URL for PDF To Image API call
string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated JPEG files
int page = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\page{0}.{1}", page, imageType);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
page++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg");
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException
{
// Prepare URL for PDF To Image API call
String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType);
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated JPEG files
JsonArray urls = json.get("urls").getAsJsonArray();
int page = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\page%s.%s", page, imageType);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Downloaded \"%s\".", localFileName));
page++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Image Results
Status code: " . $status_code . "";
echo "
";
}
else
{
// JPEG and PNG formats are single-page, so the results are multiple
$resultFiles = $json["urls"];
foreach ($resultFiles as &$resultFileUrl)
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to WEBP
Source: https://developer.pdf.co/api/pdf-to-image/webp
PDF to high quality WEBP image conversion. High quality rendering. Also works great for thumbnail generation and previews.
**Try it live:** [PDF to WEBP → API Tester](/api-tester/pdf-to-image/webp) — send a real request from your browser.
## `POST /v1/pdf/convert/to/webp`
`WEBP` is an image format invented by Google and is supported by Google Chrome and other modern browsers. It provides good quality with smaller file sizes.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| ----------------------------- | ------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `outputDataFormat` | string | *No* | - | If you require your output as `base64` format, set this to `base64` |
| `RenderTextObjects` | boolean | *No* | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](#disable-text-layer) for more information. |
| `RenderingResolution` | integer | *No* | 120 | See [Set Image Resolution](#set-image-resolution) for more information. |
| `RenderImageObjects` | boolean | *No* | `true` | Render image objects or not |
| `RenderVectorObjects` | boolean | *No* | `true` | Render vector objects or not |
| `WEBPQuality` | integer | *No* | 75 | See [profiles.WEBPQuality](#profiles-webpquality) |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Disable Text Layer
We can turn off the text layer for our render as follows:
```json theme={null}
{
"profiles": "{ 'RenderTextObjects': false }"
}
```
### Set Image Resolution
By default the screen resolution is 120 DPI. To change the rendering resolution, please use:
```json theme={null}
{
"profiles": "{ 'RenderingResolution': 300 }"
}
```
### `profiles.WEBPQuality`
To control the quality and encoding speed use the following:
```
{
"profiles": "{ 'WEBPQuality': 75 }"
}
```
### Use this parameter to set additional configurations for fine-tuning and extra options. Explore the Profiles section for more.
Profiles section for more.
```
"profiles":
{
'RenderTextObjects': false,
'RenderVectorObjects': true,
'RenderImageObjects': true
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------ |
| `urls` | array\[string] | List of URLs to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"urls": [
"https://pdf-temp-files.s3.amazonaws.com/c15b8d82e0034d01a73eac719d69349b/sample.png",
"https://pdf-temp-files.s3.amazonaws.com/152d2fe414b645e38f81a49e5dafa85b/sample.png"
],
"pageCount": 2,
"error": false,
"status": 200,
"name": "sample.png",
"duration": 1121,
"remainingCredits": 98722216,
"credits": 30
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/png' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-image/sample.pdf",
"inline": true,
"pages": "0-",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
convertPDFToImage(API_KEY, uploadedFileUrl, Password, Pages, "jpg");
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
// `imageType` should correspond to the API we want to use,
// i.e. /v1/pdf/convert/to/jpg, /v1/pdf/convert/to/png, /v1/pdf/convert/to/webp or /v1/pdf/convert/to/tiff
// we just take the last part of the path, the file extension
function convertPDFToImage(apiKey, uploadedFileUrl, password, pages, imageType) {
// Prepare URL for PDF to Image API call
var queryPath = `/v1/pdf/convert/to/${imageType}`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download generated image files
var page = 1;
data.urls.forEach((url) => {
var localFileName = `./page${page}.${imageType}`;
var file = fs.createWriteStream(localFileName);
https.get(url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated image file saved as "${localFileName}" file.`);
});
});
page++;
}, this);
}
else {
// Service reported error
console.log("convertPDFToImage(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPDFToImage(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPDFToImage(uploadedFileUrl, "jpg")
def convertPDFToImage(uploadedFileUrl, imageType):
"""Converts PDF To Image using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-image/webp
parameters = {}
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for PDF To Image API request
url = "{}/pdf/convert/to/{}".format(BASE_URL, imageType)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Download generated JPG files
part = 1
for resultFileUrl in json["urls"]:
# Download Result File
r = requests.get(resultFileUrl, stream=True)
localFileUrl = f"Page{part}.{imageType}"
if r.status_code == 200:
with open(localFileUrl, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{localFileUrl}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
part = part + 1
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
const string imageType = "jpg";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have the direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
// Get URL of uploaded file to use with later API calls
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
// Prepare URL for PDF To Image API call
string url = String.Format(@"https://api.pdf.co/v1/pdf/convert/to/{0}", imageType);
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Download generated JPEG files
int page = 1;
foreach (JToken token in json["urls"])
{
string resultFileUrl = token.ToString();
string localFileName = String.Format(@".\page{0}.{1}", page, imageType);
webClient.DownloadFile(resultFileUrl, localFileName);
Console.WriteLine("Downloaded \"{0}\".", localFileName);
page++;
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonArray;
import com.google.gson.JsonElement;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO IMAGE
PDFToImage(webClient, API_KEY, Password, Pages, uploadedFileUrl, "jpg");
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PDFToImage(OkHttpClient webClient, String apiKey, String password, String pages, String uploadedFileUrl, String imageType) throws IOException
{
// Prepare URL for PDF To Image API call
String query = String.format("https://api.pdf.co/v1/pdf/convert/to/%s", imageType);
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Download generated JPEG files
JsonArray urls = json.get("urls").getAsJsonArray();
int page = 1;
for (JsonElement element: urls)
{
String resultFileUrl = element.getAsString();
String localFileName = String.format(".\\page%s.%s", page, imageType);
downloadFile(webClient, resultFileUrl, Paths.get(localFileName).toFile());
System.out.println(String.format("Downloaded \"%s\".", localFileName));
page++;
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Image Results
Status code: " . $status_code . "";
echo "
";
}
else
{
// JPEG and PNG formats are single-page, so the results are multiple
$resultFiles = $json["urls"];
foreach ($resultFiles as &$resultFileUrl)
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to JSON
Source: https://developer.pdf.co/api/pdf-to-json/basic
Convert PDF and scanned images into JSON representation with text, fonts, images, vectors, and formatting preserved.
**Try it live:** [PDF to JSON → API Tester](/api-tester/pdf-to-json/basic) — send a real request from your browser.
## `POST /v1/pdf/convert/to/json2`
This endpoint can also be used with a specified [profile to extract image data from a PDF into your JSON output](/api/profiles/#api-profiles-save-images).You can extract hyperlinks from your PDF by using a [profile to extract only link objects](/api/profiles#extract-hyperlinks).
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the JSON file.
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description | | |
| ------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ | - | - |
| `body` | object | *No* | - | - |
| `properties` | object | *No* | - | - |
| `document` | object | *No* | - | - |
| `properties` | object | *No* | - | - |
| `pageCount` | string | Total number of pages in the document. | | |
| `pageCountWithOCRPerformed` | string | Total number of pages in the document with OCR performed. | | |
| `page` | object | Page details. | | |
| `pageCount` | integer | Number of pages in the PDF document. | | |
| `error` | boolean | Indicates whether an error occurred (`false` means success) | | |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). | | |
| `name` | string | Name of the output file | | |
| `credits` | integer | Number of credits consumed by the request | | |
| `remainingCredits` | integer | Number of credits remaining in the account | | |
| `duration` | integer | Time taken for the operation in milliseconds | | |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf",
"inline": true,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"url": "https://example.com/file1.pdf"
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/json2' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf",
"inline": true,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination JSON file name
const DestinationFile = "./result.json";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO JSON
convertPdfToJson(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToJson(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To JSON` API endpoint
var queryPath = `/v1/pdf/convert/to/json`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download JSON file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated JSON file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToJson(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToJson(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination JSON file name
DestinationFile = ".\\result.json"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToJson(uploadedFileUrl, DestinationFile)
def convertPdfToJson(uploadedFileUrl, destinationFile):
"""Converts PDF To Json using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-json/basic
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Json' API request
url = "{}/pdf/convert/to/json".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination JSON file name
const string DestinationFile = @".\result.json";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO JSON
// URL for `PDF To JSON` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/json";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated JSON file
string resultFileUrl = json["url"].ToString();
// Download JSON file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination JSON file name
final static Path DestinationFile = Paths.get(".\\result.json");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO JSON
PdfToJson(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToJson(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To JSON` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/json";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated JSON file
String resultFileUrl = json.get("url").getAsString();
// Download JSON file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To JSON Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to JSON with AI
Source: https://developer.pdf.co/api/pdf-to-json/with-ai
Convert PDF and scanned images into JSON using AI.
**Try it live:** [PDF to JSON with AI → API Tester](/api-tester/pdf-to-json/with-ai) — send a real request from your browser.
## `POST /v1/pdf/convert/to/json-meta`
This endpoint can also be used with a specified [profile to extract image data from a PDF into your JSON output](/api/profiles/#api-profiles-save-images).
What is the difference between `/pdf/convert/to/json-meta` and `/pdf/convert/to/json2`?
`/pdf/convert/to/json-meta` uses AI to detect meta styles for text objects, such as:
* paragraph style (from `h1` .. `h7` to `p` and `small`).
* meta `type` of the text object (`text`, `datetime`, `integer`, `decimal`, `currency` etc.).
* meta `subType` of the text object (`companyName`, `personName` and other AI-based meta types).
* `/json-meta` consumes more credits because it runs with AI.
* `/json-meta` is also a bit slower due to the AI process. `Async` mode is recommended for this endpoint.
Convert PDF and scanned images into JSON using AI.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | None | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file.
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| ------------------ | ------- | -------------------------------------- |
| `body` | object | Response body. |
| `pageCount` | integer | Total number of pages in the document. |
| `error` | boolean | Indicates if an error occurred. |
| `status` | integer | Status code. |
| `name` | string | Name of the job. |
| `credits` | integer | Total number of credits used. |
| `remainingCredits` | integer | Remaining number of credits. |
| `duration` | integer | Duration of the job. |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf",
"inline": true,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": {
"document": {
"pageCount": "1",
"pageCountWithOCRPerformed": "0",
"page": {
"index": "0",
"width": "595.320007324219",
"height": "841.919982910156",
"OCRWasPerformed": "False",
"row": [
{
"column": [
{
"text": {
"fontName": "Arial",
"fontSize": "24.0",
"fontStyle": "Bold",
"color": "#538DD3",
"x": "36.00",
"y": "34.44",
"width": "242.81",
"height": "24.00",
"text": "Your Company Name"
}
},
{
"text": ""
},
{
"text": ""
},
{
"text": ""
}
]
},
{
"column": [
{
"text": ""
},
{
"text": ""
},
{
"text": {
"fontName": "Arial",
"fontSize": "11.0",
"fontStyle": "Bold",
"x": "389.11",
"y": "425.83",
"width": "36.75",
"height": "11.04",
"text": "TOTAL"
}
},
{
"text": {
"fontName": "Arial",
"fontSize": "11.0",
"fontStyle": "Bold",
"x": "525.82",
"y": "425.83",
"width": "33.62",
"height": "11.04",
"text": "200.00"
}
}
]
}
]
}
}
},
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.json",
"remainingCredits": 99227903,
"credits": 28
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/json-meta' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf",
"inline": true,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination JSON file name
const DestinationFile = "./result.json";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO JSON
convertPdfToJson(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToJson(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To JSON` API endpoint
var queryPath = `/v1/pdf/convert/to/json`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download JSON file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated JSON file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToJson(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToJson(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination JSON file name
DestinationFile = ".\\result.json"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToJson(uploadedFileUrl, DestinationFile)
def convertPdfToJson(uploadedFileUrl, destinationFile):
"""Converts PDF To Json using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-json/with-ai
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Json' API request
url = "{}/pdf/convert/to/json".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination JSON file name
const string DestinationFile = @".\result.json";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO JSON
// URL for `PDF To JSON` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/json";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated JSON file
string resultFileUrl = json["url"].ToString();
// Download JSON file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated JSON file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination JSON file name
final static Path DestinationFile = Paths.get(".\\result.json");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO JSON
PdfToJson(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToJson(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To JSON` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/json";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated JSON file
String resultFileUrl = json.get("url").getAsString();
// Download JSON file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated JSON file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To JSON Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to Text
Source: https://developer.pdf.co/api/pdf-to-text/basic
Convert PDF and scanned images to text with layout preserved. This method uses OCR and reporoduces layout.
**Try it live:** [PDF to Text → API Tester](/api-tester/pdf-to-text/basic) — send a real request from your browser.
## `POST /v1/pdf/convert/to/text`
**Auto classification Of incoming documents**: Use the [Document Classifier](/api/document-classifier) endpoint to automatically sort/detect the class of the document based on keywords-based rules. For example, you can define rules to find which vendor provided the document to find which template to apply accordingly.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | eng | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"inline": true,
"async": false
}'
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": " Your Company Name \r\n Your Address \r\n City, State Zip \r\n Invoice No. 123456 \r\n Invoice Date 01/01/2016 \r\n Client Name \r\n Address \r\n City, State Zip \r\n\r\n Notes \r\n\r\n\r\n Item Quantity Price Total \r\n Item 1 1 40.00 40.00 \r\n Item 2 2 30.00 60.00 \r\n Item 3 3 20.00 60.00 \r\n Item 4 4 10.00 40.00 \r\n TOTAL 200.00\r\n",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.txt",
"remainingCredits": 99032333,
"credits": 21
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"inline": true,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination TXT file name
const DestinationFile = "./result.txt";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO TEXT
convertPdfToText(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToText(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To Text` API endpoint
var queryPath = `/v1/pdf/convert/to/text`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download TXT file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated TXT file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToText(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToText(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination TXT file name
DestinationFile = ".\\result.txt"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToText(uploadedFileUrl, DestinationFile)
def convertPdfToText(uploadedFileUrl, destinationFile):
"""Converts PDF To Text using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-text/basic
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Text' API request
url = "{}/pdf/convert/to/text".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination TXT file name
const string DestinationFile = @".\result.txt";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO TXT
// URL for `PDF To TXT` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/text";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated TXT file
string resultFileUrl = json["url"].ToString();
// Download TXT file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated TXT file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination TXT file name
final static Path DestinationFile = Paths.get(".\\result.txt");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO TXT
PdfToText(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToText(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To TXT` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/text";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated TXT file
String resultFileUrl = json.get("url").getAsString();
// Download TXT file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated TXT file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Text Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to Text (Simple)
Source: https://developer.pdf.co/api/pdf-to-text/simple
Extract plain text from PDF documents using a fast, low-credit method without OCR, layout analysis, or profile-based fine-tuning.
**Try it live:** [PDF to Text (Simple) → API Tester](/api-tester/pdf-to-text/simple) — send a real request from your browser.
## `POST /v1/pdf/convert/to/text-simple`
**Auto classification Of incoming documents**: Use the [Document Classifier](/api/document-classifier) endpoint to automatically sort/detect the class of the document based on keywords-based rules. For example, you can define rules to find which vendor provided the document to find which template to apply accordingly.
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------- | ------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string / integer | Status of the API response. Returns the string `success` on success. On failure it is either a numeric code such as `400` for endpoint-level errors, or the string `error` for authentication and routing failures, and the numeric code is also returned in the `errorCode` field. For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"inline": true,
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": "Your Company Name \r\nYour Address \r\nCity, State Zip \r\nInvoice No. 123456 \r\nInvoice Date 01/01/2016 \r\nClient Name \r\nAddress \r\nCity, State Zip \r\nNotes \r\nItem Quantity Price Total \r\nItem 1 1 40.00 40.00 \r\nItem 2 2 30.00 60.00 \r\nItem 3 3 20.00 60.00 \r\nItem 4 4 10.00 40.00 \r\nTOTAL 200.00 \r\n",
"pageCount": 1,
"error": false,
"status": "success",
"name": "sample.txt",
"remainingCredits": 99885491,
"credits": 4
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/text-simple' \
--header 'Content-Type: application/json' \
--header 'x-api-key: *******************' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"inline": true,
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination TXT file name
const DestinationFile = "./result.txt";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO TEXT
convertPdfToText(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(localFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(localFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToText(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To Text (Simple)` API endpoint
var queryPath = `/v1/pdf/convert/to/text-simple`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download TXT file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated TXT file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToText(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToText(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination TXT file name
DestinationFile = ".\\result.txt"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToText(uploadedFileUrl, DestinationFile)
def convertPdfToText(uploadedFileUrl, destinationFile):
"""Converts PDF To Text using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-text/simple
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To Text (Simple)' API request
url = "{}/pdf/convert/to/text-simple".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination TXT file name
const string DestinationFile = @".\result.txt";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO TXT
// URL for `PDF To Text (Simple)` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/text-simple";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated TXT file
string resultFileUrl = json["url"].ToString();
// Download TXT file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated TXT file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination TXT file name
final static Path DestinationFile = Paths.get(".\\result.txt");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO TXT
PdfToText(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToText(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To Text (Simple)` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/text-simple";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated TXT file
String resultFileUrl = json.get("url").getAsString();
// Download TXT file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated TXT file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To Text Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# PDF to XML
Source: https://developer.pdf.co/api/pdf-to-xml
Convert PDF to XML with information about text value, tables, fonts, images, objects positions.
**Try it live:** [PDF to XML → API Tester](/api-tester/pdf-to-xml) — send a real request from your browser.
## `POST /v1/pdf/convert/to/xml`
## Attributes
Attributes are case-sensitive and should be inside JSON for POST request. for example: `{ "url": "https://example.com/file1.pdf" }`
| Attribute | Type | Required | Default | Description |
| -------------------------------------- | ----------------------------- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `url` | string | *Yes* | - | URL to the source file [`url` attribute](/api/url-input-and-request-limits#supported-file-sources) |
| `callback` | string | *No* | - | The callback URL (or Webhook) used to receive the POST data. see [Webhooks & Callbacks](/api/webhooks). This is only applicable when `async` is set to `true`. |
| `httpusername` | string | *No* | - | HTTP auth user name if required to access source URL. |
| `httppassword` | string | *No* | - | HTTP auth password if required to access source URL. |
| `pages` | string | *No* | all pages | Specify page indices as comma-separated values or ranges to process (e.g. "0, 1, 2-" or "1, 2, 3-7"). The first-page index is 0. Use "!" before a number for inverted page numbers (e.g. "!0" for the last page). If not specified, the default configuration processes all pages. The input must be in string format. |
| `unwrap` | boolean | *No* | `false` | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. |
| `rect` | string | *No* | - | Defines coordinates for extraction. Use`PDF Edit Add Helper`to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. |
| `lang` | string | *No* | `eng` | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). |
| `inline` | boolean | *No* | `false` | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. |
| `lineGrouping` | string | *No* | - | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](#line-grouping-options). |
| `password` | string | *No* | - | Password for the PDF file. |
| `async` | boolean | *No* | `false` | Set `async` to `true` for long processes to run in the background, API will then return a `jobId` which you can use with the [Background Job Check endpoint](/api/job-check). Also see [Webhooks & Callbacks](/api/webhooks) |
| `name` | string | *No* | - | File name for the generated output, the input must be in string format. |
| `expiration` | integer | *No* | `60` | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). |
| `profiles` | object | *No* | - | See [Profiles](/api/profiles) for more information. |
| `OCRMode` | string | *No* | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [OCR Extraction Modes](/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | *No* | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | *No* | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | *No* | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | *No* | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | *No* | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | *No* | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | *No* | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | *No* | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | *No* | \["1.4"] | Adds a gamma correction filter. |
| `OCRAutoModeMinExistingTextLength` | integer | *No* | `8` | The minimum number of characters a page must have to skip OCR. If a page has fewer, OCR will run. For example, if set to 8, OCR is skipped on pages with more than 8 characters. |
| `SaveVectors` | boolean | *No* | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. |
| `SaveImages` | string | *No* | `None` | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI). |
| `ConsiderFontSizes` | boolean | *No* | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. |
| `ExtractionArea` | array\[numbe] | *No* | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. |
| `ExtractShadowLikeText` | boolean | *No* | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect. |
| `DataEncryptionAlgorithm` | string | *No* | - | Controls the encryption algorithm used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | *No* | - | Controls the encryption key used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | *No* | - | Controls the encryption IV used for data encryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | *No* | - | Controls the decryption algorithm used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | *No* | - | Controls the decryption key used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | *No* | - | Controls the decryption IV used for data decryption. See [User-Controlled Encryption](/knowledgebase/user-controlled-encryption) for more information. |
You can use [profiles](/api/profiles#converting-pdfs) to control the convert process and output of the CSV file.
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
### Line Grouping Options
* `"1"`: GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
* `"2"`: GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can't be grouped. Useful for columnar data where content in each column might span multiple lines.
* `"3"`: JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.
## Query parameters
*No query parameters accepted.*
## Responses
| Parameter | Type | Description |
| --------------------- | ------- | ------------------------------------------------------------------------------------------------------------------ |
| `url` | string | Direct URL to the final PDF file stored in S3. |
| `outputLinkValidTill` | string | Timestamp indicating when the output link will expire |
| `pageCount` | integer | Number of pages in the PDF document. |
| `error` | boolean | Indicates whether an error occurred (`false` means success) |
| `status` | string | Status code of the request (200, 404, 500, etc.). For more information, see [Response Codes](/api/response-codes). |
| `name` | string | Name of the output file |
| `credits` | integer | Number of credits consumed by the request |
| `remainingCredits` | integer | Number of credits remaining in the account |
| `duration` | integer | Time taken for the operation in milliseconds |
## `Example` Payload
To see the request size limits, please refer to the [Request Size Limits](/api/url-input-and-request-limits#pdf-co-request-size).
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-xml/sample.pdf",
"async": false
}
```
## `Example` Response
To see the main response codes, please refer to the [Response Codes](/api/response-codes) page.
```json theme={null}
{
"body": "\r\n\r\n \r\n \r\n \r\n Your Company Name\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Your Address\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n City, State Zip\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Invoice No. 123456\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Invoice Date 01/01/2016\r\n \r\n \r\n \r\n \r\n Client Name\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Address\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n City, State Zip\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Notes\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n Item\r\n \r\n \r\n Quantity\r\n \r\n \r\n Price\r\n \r\n \r\n Total\r\n \r\n \r\n \r\n \r\n Item 1\r\n \r\n \r\n 1\r\n \r\n \r\n 40.00\r\n \r\n \r\n 40.00\r\n \r\n \r\n \r\n \r\n Item 2\r\n \r\n \r\n 2\r\n \r\n \r\n 30.00\r\n \r\n \r\n 60.00\r\n \r\n \r\n \r\n \r\n Item 3\r\n \r\n \r\n 3\r\n \r\n \r\n 20.00\r\n \r\n \r\n 60.00\r\n \r\n \r\n \r\n \r\n Item 4\r\n \r\n \r\n 4\r\n \r\n \r\n 10.00\r\n \r\n \r\n 40.00\r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n \r\n TOTAL\r\n \r\n \r\n 200.00\r\n \r\n \r\n \r\n",
"pageCount": 1,
"error": false,
"status": 200,
"name": "sample.xml",
"remainingCredits": 60563
}
```
**Inconsistent URL Encoding in cURL Output:** When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (`&`) may appear as `\u0026` in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you're parsing the response programmatically, your JSON parser will handle this conversion automatically.
## Code Samples
```bash theme={null}
curl --location --request POST 'https://api.pdf.co/v1/pdf/convert/to/xml' \
--header 'x-api-key: *******************' \
--header 'Content-Type: application/json' \
--data-raw '{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-xml/sample.pdf",
"async": false
}'
```
```javascript theme={null}
var https = require("https");
var path = require("path");
var fs = require("fs");
// `request` module is required for file upload.
// Use "npm install request" command to install.
var request = require("request");
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const API_KEY = "***********************************";
// Source PDF file
const SourceFile = "./sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const Pages = "";
// PDF document password. Leave empty for unprotected documents.
const Password = "";
// Destination XML file name
const DestinationFile = "./result.xml";
// 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
getPresignedUrl(API_KEY, SourceFile)
.then(([uploadUrl, uploadedFileUrl]) => {
// 2. UPLOAD THE FILE TO CLOUD.
uploadFile(API_KEY, SourceFile, uploadUrl)
.then(() => {
// 3. CONVERT UPLOADED PDF FILE TO XML
convertPdfToXml(API_KEY, uploadedFileUrl, Password, Pages, DestinationFile);
})
.catch(e => {
console.log(e);
});
})
.catch(e => {
console.log(e);
});
function getPresignedUrl(apiKey, localFile) {
return new Promise(resolve => {
// Prepare request to `Get Presigned URL` API endpoint
let queryPath = `/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=${path.basename(SourceFile)}`;
let reqOptions = {
host: "api.pdf.co",
path: encodeURI(queryPath),
headers: { "x-api-key": API_KEY }
};
// Send request
https.get(reqOptions, (response) => {
response.on("data", (d) => {
let data = JSON.parse(d);
if (data.error == false) {
// Return presigned url we received
resolve([data.presignedUrl, data.url]);
}
else {
// Service reported error
console.log("getPresignedUrl(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("getPresignedUrl(): " + e);
});
});
}
function uploadFile(apiKey, localFile, uploadUrl) {
return new Promise(resolve => {
fs.readFile(SourceFile, (err, data) => {
request({
method: "PUT",
url: uploadUrl,
body: data,
headers: {
"Content-Type": "application/octet-stream"
}
}, (err, res, body) => {
if (!err) {
resolve();
}
else {
console.log("uploadFile() request error: " + e);
}
});
});
});
}
function convertPdfToXml(apiKey, uploadedFileUrl, password, pages, destinationFile) {
// Prepare request to `PDF To XML` API endpoint
var queryPath = `/v1/pdf/convert/to/xml`;
// JSON payload for api request
var jsonPayload = JSON.stringify({
name: path.basename(destinationFile), password: password, pages: pages, url: uploadedFileUrl
});
var reqOptions = {
host: "api.pdf.co",
method: "POST",
path: queryPath,
headers: {
"x-api-key": apiKey,
"Content-Type": "application/json",
"Content-Length": Buffer.byteLength(jsonPayload, 'utf8')
}
};
// Send request
var postRequest = https.request(reqOptions, (response) => {
response.on("data", (d) => {
response.setEncoding("utf8");
// Parse JSON response
let data = JSON.parse(d);
if (data.error == false) {
// Download XML file
var file = fs.createWriteStream(destinationFile);
https.get(data.url, (response2) => {
response2.pipe(file)
.on("close", () => {
console.log(`Generated XML file saved as "${destinationFile}" file.`);
});
});
}
else {
// Service reported error
console.log("convertPdfToXml(): " + data.message);
}
});
})
.on("error", (e) => {
// Request error
console.log("convertPdfToXml(): " + e);
});
// Write request data
postRequest.write(jsonPayload);
postRequest.end();
}
```
```python theme={null}
import os
import requests # pip install requests
# The authentication key (API Key).
# Get your own by registering at https://app.pdf.co
API_KEY = "******************************************"
# Base URL for PDF.co Web API requests
BASE_URL = "https://api.pdf.co/v1"
# Source PDF file
SourceFile = ".\\sample.pdf"
# Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
Pages = ""
# PDF document password. Leave empty for unprotected documents.
Password = ""
# Destination XML file name
DestinationFile = ".\\result.xml"
def main(args = None):
uploadedFileUrl = uploadFile(SourceFile)
if (uploadedFileUrl != None):
convertPdfToXml(uploadedFileUrl, DestinationFile)
def convertPdfToXml(uploadedFileUrl, destinationFile):
"""Converts PDF To XML using PDF.co Web API"""
# Prepare requests params as JSON
# See documentation: https://developer.pdf.co/api/pdf-to-xml
parameters = {}
parameters["name"] = os.path.basename(destinationFile)
parameters["password"] = Password
parameters["pages"] = Pages
parameters["url"] = uploadedFileUrl
# Prepare URL for 'PDF To XML' API request
url = "{}/pdf/convert/to/xml".format(BASE_URL)
# Execute request and get response as JSON
response = requests.post(url, data=parameters, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# Get URL of result file
resultFileUrl = json["url"]
# Download result file
r = requests.get(resultFileUrl, stream=True)
if (r.status_code == 200):
with open(destinationFile, 'wb') as file:
for chunk in r:
file.write(chunk)
print(f"Result file saved as \"{destinationFile}\" file.")
else:
print(f"Request error: {response.status_code} {response.reason}")
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
def uploadFile(fileName):
"""Uploads file to the cloud"""
# 1. RETRIEVE PRESIGNED URL TO UPLOAD FILE.
# Prepare URL for 'Get Presigned URL' API request
url = "{}/file/upload/get-presigned-url?contenttype=application/octet-stream&name={}".format(
BASE_URL, os.path.basename(fileName))
# Execute request and get response as JSON
response = requests.get(url, headers={ "x-api-key": API_KEY })
if (response.status_code == 200):
json = response.json()
if json["error"] == False:
# URL to use for file upload
uploadUrl = json["presignedUrl"]
# URL for future reference
uploadedFileUrl = json["url"]
# 2. UPLOAD FILE TO CLOUD.
with open(fileName, 'rb') as file:
requests.put(uploadUrl, data=file, headers={ "x-api-key": API_KEY, "content-type": "application/octet-stream" })
return uploadedFileUrl
else:
# Show service reported error
print(json["message"])
else:
print(f"Request error: {response.status_code} {response.reason}")
return None
if __name__ == '__main__':
main()
```
```csharp theme={null}
using System;
using System.Collections.Generic;
using System.IO;
using System.Net;
using Newtonsoft.Json;
using Newtonsoft.Json.Linq;
namespace PDFcoApiExample
{
class Program
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
const String API_KEY = "***********************************";
// Source PDF file
const string SourceFile = @".\sample.pdf";
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
const string Pages = "";
// PDF document password. Leave empty for unprotected documents.
const string Password = "";
// Destination XML file name
const string DestinationFile = @".\result.xml";
static void Main(string[] args)
{
// Create standard .NET web client instance
WebClient webClient = new WebClient();
// Set API Key
webClient.Headers.Add("x-api-key", API_KEY);
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
string query = Uri.EscapeUriString(string.Format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name={0}",
Path.GetFileName(SourceFile)));
try
{
// Execute request
string response = webClient.DownloadString(query);
// Parse JSON response
JObject json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL to use for the file upload
string uploadUrl = json["presignedUrl"].ToString();
string uploadedFileUrl = json["url"].ToString();
// 2. UPLOAD THE FILE TO CLOUD.
webClient.Headers.Add("content-type", "application/octet-stream");
webClient.UploadFile(uploadUrl, "PUT", SourceFile); // You can use UploadData() instead if your file is byte[] or Stream
webClient.Headers.Remove("content-type");
// 3. CONVERT UPLOADED PDF FILE TO XML
// URL for `PDF To XML` API call
var url = "https://api.pdf.co/v1/pdf/convert/to/xml";
// Prepare requests params as JSON
Dictionary parameters = new Dictionary();
parameters.Add("name", Path.GetFileName(DestinationFile));
parameters.Add("password", Password);
parameters.Add("pages", Pages);
parameters.Add("url", uploadedFileUrl);
// Convert dictionary of params to JSON
string jsonPayload = JsonConvert.SerializeObject(parameters);
// Execute POST request with JSON payload
response = webClient.UploadString(url, jsonPayload);
// Parse JSON response
json = JObject.Parse(response);
if (json["error"].ToObject() == false)
{
// Get URL of generated XML file
string resultFileUrl = json["url"].ToString();
// Download XML file
webClient.DownloadFile(resultFileUrl, DestinationFile);
Console.WriteLine("Generated XML file saved as \"{0}\" file.", DestinationFile);
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
else
{
Console.WriteLine(json["message"].ToString());
}
}
catch (WebException e)
{
Console.WriteLine(e.ToString());
}
webClient.Dispose();
Console.WriteLine();
Console.WriteLine("Press any key...");
Console.ReadKey();
}
}
}
```
```java theme={null}
package com.company;
import com.google.gson.JsonObject;
import com.google.gson.JsonParser;
import okhttp3.*;
import java.io.*;
import java.net.*;
import java.nio.file.Path;
import java.nio.file.Paths;
public class Main
{
// The authentication key (API Key).
// Get your own by registering at https://app.pdf.co
final static String API_KEY = "***********************************";
// Source PDF file
final static Path SourceFile = Paths.get(".\\sample.pdf");
// Comma-separated list of page indices (or ranges) to process. Leave empty for all pages. Example: '0,2-5,7-'.
final static String Pages = "";
// PDF document password. Leave empty for unprotected documents.
final static String Password = "";
// Destination XML file name
final static Path DestinationFile = Paths.get(".\\result.xml");
public static void main(String[] args) throws IOException
{
// Create HTTP client instance
OkHttpClient webClient = new OkHttpClient();
// 1. RETRIEVE THE PRESIGNED URL TO UPLOAD THE FILE.
// * If you already have a direct file URL, skip to the step 3.
// Prepare URL for `Get Presigned URL` API call
String query = String.format(
"https://api.pdf.co/v1/file/upload/get-presigned-url?contenttype=application/octet-stream&name=%s",
SourceFile.getFileName());
// Prepare request
Request request = new Request.Builder()
.url(query)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL to use for the file upload
String uploadUrl = json.get("presignedUrl").getAsString();
// Get URL of uploaded file to use with later API calls
String uploadedFileUrl = json.get("url").getAsString();
// 2. UPLOAD THE FILE TO CLOUD.
if (uploadFile(webClient, API_KEY, uploadUrl, SourceFile))
{
// 3. CONVERT UPLOADED PDF FILE TO XML
PdfToXml(webClient, API_KEY, DestinationFile, Password, Pages, uploadedFileUrl);
}
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static void PdfToXml(OkHttpClient webClient, String apiKey, Path destinationFile,
String password, String pages, String uploadedFileUrl) throws IOException
{
// Prepare URL for `PDF To XML` API call
String query = "https://api.pdf.co/v1/pdf/convert/to/xml";
// Make correctly escaped (encoded) URL
URL url = null;
try
{
url = new URI(null, query, null).toURL();
}
catch (URISyntaxException e)
{
e.printStackTrace();
}
// Create JSON payload
String jsonPayload = String.format("{\"name\": \"%s\", \"password\": \"%s\", \"pages\": \"%s\", \"url\": \"%s\"}",
destinationFile.getFileName(),
password,
pages,
uploadedFileUrl);
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/json"), jsonPayload);
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", API_KEY) // (!) Set API Key
.addHeader("Content-Type", "application/json")
.post(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
if (response.code() == 200)
{
// Parse JSON response
JsonObject json = new JsonParser().parse(response.body().string()).getAsJsonObject();
boolean error = json.get("error").getAsBoolean();
if (!error)
{
// Get URL of generated XML file
String resultFileUrl = json.get("url").getAsString();
// Download XML file
downloadFile(webClient, resultFileUrl, destinationFile.toFile());
System.out.printf("Generated XML file saved as \"%s\" file.", destinationFile.toString());
}
else
{
// Display service reported error
System.out.println(json.get("message").getAsString());
}
}
else
{
// Display request error
System.out.println(response.code() + " " + response.message());
}
}
public static boolean uploadFile(OkHttpClient webClient, String apiKey, String url, Path sourceFile) throws IOException
{
// Prepare request body
RequestBody body = RequestBody.create(MediaType.parse("application/octet-stream"), sourceFile.toFile());
// Prepare request
Request request = new Request.Builder()
.url(url)
.addHeader("x-api-key", apiKey) // (!) Set API Key
.addHeader("content-type", "application/octet-stream")
.put(body)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
return (response.code() == 200);
}
public static void downloadFile(OkHttpClient webClient, String url, File destinationFile) throws IOException
{
// Prepare request
Request request = new Request.Builder()
.url(url)
.build();
// Execute request
Response response = webClient.newCall(request).execute();
byte[] fileBytes = response.body().bytes();
// Save downloaded bytes to file
OutputStream output = new FileOutputStream(destinationFile);
output.write(fileBytes);
output.flush();
output.close();
response.close();
}
}
```
```php theme={null}
PDF To XML Extraction Results
Status code: " . $status_code . "";
echo "
";
}
}
else
{
// Display CURL error
echo "Error: " . curl_error($curl);
}
// Cleanup
curl_close($curl);
}
?>
```
# Postman
Source: https://developer.pdf.co/api/postman
Import the PDF.co Postman Collection to explore and test every API endpoint.
[Postman](https://www.postman.com/) is the collaboration platform for API development, used by 10 million developers and 500,000 companies worldwide. The Postman API Platform simplifies each step of building an API, and streamlines collaboration, so you can create better APIs — faster.
## Getting Started with Postman & PDF.co
We have created **PDF.co Collection** for **Postman** that you can import into **Postman** and explore **PDF.co API** and functions right away.
Before beginning the **Postman Collection** tutorial, please do the following:
* Download and install the [Postman](https://www.postman.com/) app.
* Download the [PDF.co Postman Collection](https://pdfco-docs-mintlify.s3.ap-southeast-2.amazonaws.com/PDF.co+API+v.1.00.postman_collection.json).
* Import this `.JSON` file into Postman and set up your API key for use with tests. Make sure that you have your API Key ready.
# Profiles
Source: https://developer.pdf.co/api/profiles
This page describes the `profiles` parameter that can be used with your API calls.
Profiles are used to to set extra options for common API calls and are sometimes distinct to a particular API.
Profiles are embedded with a `JSON` type of notation along with the `profiles` object for your API calls, for example:
Please note that the value for the `profiles` field in the code snippets must be enclosed in quotes (`"`), making it a complete string. For example: `{ "profiles": "{'TrimSpaces':true, 'PreserveFormattingOnTextExtraction': true}"}`
## Sample Code
```json theme={null}
{
"profiles": "{'TrimSpaces':true, 'PreserveFormattingOnTextExtraction': true}"
}
```
```
profiles = '"TrimSpaces": "True", "PreserveFormattingOnTextExtraction": "True" '
```
```json theme={null}
{
"profiles": "'TrimSpaces': 'True' , 'PreserveFormattingOnTextExtraction': 'True'"
}
```
```
String profiles = "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }";
```
```
const Profiles = "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }";
```
```
$Profiles = '{ "TrimSpaces": "True", "PreserveFormattingOnTextExtraction": "True" }'
```
```json theme={null}
{
"profiles": "'TrimSpaces': 'True' , 'PreserveFormattingOnTextExtraction': 'True'"
}
```
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-json/sample.pdf",
"inline": true,
"profiles": "{ 'TrimSpaces': 'True', 'PreserveFormattingOnTextExtraction': 'True' }"
}
```
***
## Generic Profile Options
The following `profiles` options are not specific to any one particular endpoint.
### Standard Parameters
The `std_params` within the `profiles` parameter enables the definition of regular API parameters in a `JSON` format. This `std_params` feature is designed to simplify the process of passing standard parameters and additional options in the `profiles` parameter for PDF.co API requests.
When using [Standard Parameters](#standard-parameters) webhooks can be utilized by setting the `callback` object with the URL of your choice. However, is is simpler to set the `callback` object directly - see [Webhooks & Callbacks](/api/webhooks) for more.
When `std_params` are used in the `profiles` parameter, if a parameter is duplicated within both `std_params` and outside profiles, the value specified in `std_params` will overwrite the duplicate value. Therefore if you define a callback object in `std_params` then it will overwrite any value you may have defined via [the basic callback object](/api/webhooks)!
#### `std_params` Structure
* **Description**: Contains key-value pairs of standard parameters that will be used across PDF.co API requests.
* **Type**: `JSON` Object (passed as a string)
* **Example**:
```json theme={null}
{
"profiles": "{'std_params': {'callback': 'webhook_url'}}"
}
```
#### Practical Application
Using the `std_params` profile, you can define a set of standard parameters and configurations that will be consistently applied across your PDF.co API requests. This approach is particularly beneficial when using automation platforms like [Zapier](/integrations/zapier/), [Make](/integrations/make/), and others, where the number of parameters you can pass directly is limited.
#### Complete Request Example
Here is a complete example illustrating the use of the `std_params` profile with other parameters:
/pdf/convert/to/text
```json theme={null}
{
"url": "https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-text/sample.pdf",
"inline": true,
"profiles": "{'std_params': {'callback': 'webhook_url', 'async': true}, 'ExtractShadowLikeText': false, 'ExtractColumnByColumn': true, 'OCRMode': 'Auto'}}",
"TrimSpaces": true,
"PreserveFormattingOnTextExtraction": true
}
```
### Output as Base64
If you require your output as `base64` use the following:
```json theme={null}
{
"profiles": "{ 'outputDataFormat': 'base64' }"
}
```
This output data format is supported by endpoints that generate binary files - **PDF** and images. The output is accessible via a generated link and the file under the link is in a base64-encoded text format.
### Converting PDFs
There are a variety of `profiles` options which can be set when converting from **PDF** to other documents. These `profiles` control how to extract the information from the source **PDF** file.
These options apply to the following endpoints:
* /pdf/convert/to/csv
* /pdf/convert/to/xml
* /pdf/convert/to/json
* /pdf/convert/to/json2
* /pdf/convert/to/xls
* /pdf/convert/to/xlsx
#### Convert Vectors
You can choose whether the conversion process should convert vectors or not as follows:
```json theme={null}
{
"profiles": "{ 'SaveVectors': true }"
}
```
#### Save Images
This `profiles` parameter includes the `SaveImages` property that extracts individual images in a regular **PDF**.
```json theme={null}
{
"profiles": "{ 'SaveImages': 'Embed' }"
}
```
#### Consider Font Size
This `profiles` parameter allows you to seperate header and body text based on font size.
```json theme={null}
{
"profiles": "{ 'ConsiderFontSizes': true }"
}
```
#### Set the Extraction Area
Extract text in a specific area by defining the extraction area - set with points in the format `[x, y, width, height]`.
```json theme={null}
{
"profiles": "{ 'ExtractionArea': [171.0,69.0,249.75,71.25] }"
}
```
#### Extract Hyperlinks
Extract hyperlinks (URLs) from a PDF document by using the `OutputStructure` and `OutputTransformation` profile options. This returns only the link objects found in the PDF.
```json theme={null}
{
"profiles": "{ 'OutputStructure': 'OnlyLinks', 'OutputTransformation': '$..text' }"
}
```
* `OutputStructure`: Set to `OnlyLinks` to restrict the output to hyperlink elements only.
* `OutputTransformation`: Set to `$..text` (a JSONPath expression) to extract just the link text/URL values from the result.
#### Extracting Invisible Text
When dealing with **PDF** documents, sometimes there may be unwanted invisible text that makes it difficult to extract the desired content accurately. This could be due to various reasons such as the original document being scanned or saved with a low-quality setting. In such cases, it is important to remove the unwanted invisible text to ensure accurate extraction of the desired content.
```json theme={null}
{
"profiles": "{ 'ExtractInvisibleText': false, 'ExtractShadowLikeText': false, 'OCRMode': 'Auto' }"
}
```
### OCR (Optical Character Recognition) Mode Options
The following values can be configured for OCR mode:
| OCR Mode | Description |
| ------------------------------------------ | ----------------------------------------------------------------------------- |
| `Auto` **(default)** | Automatically determines the optimal OCR settings based on the input. |
| `AutoRepairFonts` | Automatically repairs fonts in text extracted from images or other documents. |
| `TextFromImagesAndFonts` | Extracts text from images and fonts from documents. |
| `TextFromImagesAndRepairedFonts` | Extracts text from images and repaired fonts from documents. |
| `TextFromImagesAndVectorsAndFonts` | Extracts text, vectors, and fonts from images and documents. |
| `TextFromImagesAndVectorsAndRepairedFonts` | Extracts text, vectors, and repaired fonts from images and documents. |
| `TextFromImagesAndVectorsOnly` | Extracts text and vectors from images only. |
| `TextFromImagesOnly` | Extracts text from images only. |
| `TextFromRepairedFontsOnly` | Extracts text from documents with repaired fonts only. |
| `TextFromVectorsAndFonts` | Extracts text and fonts from documents with vectors. |
| `TextFromVectorsAndRepairedFonts` | Extracts text and repaired fonts from documents with vectors. |
| `TextFromVectorsOnly` | Extracts text from documents with vectors only. |
```json theme={null}
{
"profiles": "{ 'OCRMode': 'TextFromImagesAndVectorsAndRepairedFonts' }"
}
```
### OCR (Optical Character Recognition) Resolution
OCR resolution can be set from `72` to `1200` DPI. The default value is `300` DPI. The higher the resolution, the better the OCR results. However, higher resolution also means longer processing times.
```json theme={null}
{
"profiles": "{ 'OCRResolution': 300 }"
}
```
#### Extracting Text from Colored Background
If you can’t extract text with a colored background, please add the Grayscale filter to the `profiles` as follows:
```json theme={null}
{
"profiles": "{ 'OCRImagePreprocessingFilters.AddGrayscale()': [] }"
}
```
#### Considering the Font Color on Tables
Sometimes the data which OCR must extract from a table might have colored text which is difficult to extract. OCR results can be improved with the following:
```json theme={null}
{
"profiles": "{
'LineGroupingMode': 'JoinOrphanedRows',
'ConsiderFontColors': true,
'DetectNewColumnBySpacesRatio': '1.1',
'AutoAlignColumnsToHeader': false,
'OCRImagePreprocessingFilters.AddGammaCorrection()': [ '1.4' ]
}"
}
```
#### Setting the Rotation Angle
Normally OCR detects **PDF** rotation and extracts text properly. But in some cases a **PDF** is constructed in such a way that a page is not rotated and instead text is drawn vertically, OCR does not detect page rotation automatically. In such scenarios we can use following profile setting.
```json theme={null}
{
"profiles": "{ 'RotationAngle': 2 }"
}
```
* `0` no rotation
* `1` 90 degrees
* `2` 180 degrees
* `3` 270 degrees
# Response Codes
Source: https://developer.pdf.co/api/response-codes
Reference list of HTTP status codes and PDF.co-specific error codes returned by the API, with the meaning of each code.
The full set of response codes from the **PDF.co** API are as follows:
| Code | Description |
| ----- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `200` | Success. |
| `204` | No content. The server successfully processed the request, and is not returning any content. |
| `400` | Bad request. Typically due to bad input parameters or unreachable input URLs (e.g., access restrictions like login or password). |
| `401` | Unauthorized. Authentication is required and has failed or has not yet been provided. |
| `402` | Not enough credits. |
| `403` | Access forbidden for input URL. |
| `404` | The requested resource could not be found. |
| `408` | The server timed out waiting for the request. |
| `414` | The URI provided was too long for the server to process. |
| `415` | The request entity has a media type not supported by the server or resource. |
| `429` | Too many requests in a given time period. |
| `441` | Invalid Password. Password protected document. |
| `442` | Input document is damaged or of incorrect type. |
| `443` | Permissions. The operation is prohibited by document security settings. You can turn off this check by setting the `profiles` param to `{CheckPermissions: false}`. **Important:** only use this if you are the owner or have legal permission. |
| `444` | Profiles parsing error. Please ensure that the configuration is supported. See `/profiles` samples. |
| `445` | Timeout error. For large documents, use asynchronous mode (`async=true`) and check status via `/job/check`. For many-page files, use the `pages` parameter. |
| `446` | Some files required for conversion are missing. |
| `447` | Invalid template. |
| `448` | Invalid URL or HTML. Ensure the provided URL is valid and accessible. |
| `449` | Invalid index range. Page index is out of range. Use `/pdf/info` to get page count. First page is `0`. |
| `450` | Invalid page range specified. |
| `452` | Invalid URL. |
| `454` | Invalid parameters. |
| `455` | Failed to send email. |
| `456` | Invalid color. Should be a valid name (e.g., "Red") or hex string (e.g., "#CCBBAA" or "CCBBAA"). |
| `457` | SMTP server blocked. |
| `466` | Invalid base64 image. |
| `490` | Incorrect result data. |
| `500` | Something went wrong. Please try again or contact support. |
| `501` | Not implemented. The server was acting as a gateway or proxy and received an invalid response. |
| `502` | Bad gateway. The server was acting as a gateway or proxy and received an invalid response. |
| `503` | Service unavailable. Server is overloaded or under maintenance. |
| `504` | Gateway timeout. Server didn’t get a timely response from upstream. |
| `505` | HTTP version not supported. |
# URL Input and Request Limits
Source: https://developer.pdf.co/api/url-input-and-request-limits
Supported URL sources, TLS requirements, request size limits, and the cache prefix option for reusable file inputs in PDF.co API calls.
## Supported File Sources
The API supports TLS 1.2 and 1.3 for secure connections. Earlier versions such as TLS 1.0 and 1.1 are deprecated and should be avoided.
**Tip:** For re-usable files (e.g., PDF templates), use `cache:` before the URL\
(e.g., `cache:https://example.com/file1.pdf`) to reduce repeated downloads and avoid\
errors like `Access Denied or Too Many Requests`. It stores a local copy of the file after the first download, so it does not need to be fetched again from the original server.
The **PDF.co API** supports publicly accessible links from any source, including [Google Drive](https://drive.google.com), [Dropbox](https://dropbox.com), and [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files).
File inputs are accepted only via URLs and not through direct uploads. If your file is stored locally or not publicly accessible, you must upload it using the
[File Upload](/api/file-upload) endpoints to get a publicly accessible URL.
**For data security**, you have the option to **encrypt output files** and **decrypt input files**.
Learn more about [user-controlled data encryption](/knowledgebase/user-controlled-encryption).
## PDF.co Request size
API requests do not support request sizes of more than `4` megabytes in size. Please ensure that request sizes do not exceed this limit.
# Webhook and Callbacks
Source: https://developer.pdf.co/api/webhooks
Configure callback URLs and webhooks to receive PDF.co job completion events, including timeout behavior, retry logic, and Basic Authentication.
Webhooks and callbacks are both mechanisms used to enable communication between different systems or components in software development, often for the purpose of executing code in response to specific events. However, they operate in different contexts and are used in different ways.
## What are Callbacks?
Callbacks are functions that are passed as arguments to other functions or methods, and they are invoked (or "called back") after the completion of a task or operation. Callbacks are a common pattern in programming, particularly in *asynchronous* operations, where you want to execute a piece of code after a certain task is done without blocking the main execution thread.
## What are Webhooks?
Webhooks are a type of callback that operates over the web, allowing one system to send real-time data to another system as soon as an event occurs. Unlike traditional callbacks, which are typically defined within the same codebase or application, webhooks are used for communication between different applications or services over the internet.
## PDF.co & Webhooks
The webhook or callback endpoint must respond immediately within a few seconds. If the endpoint does not respond within **20-30 seconds**, the PDF.co API considers the attempt failed and will retry the callback up to **3 times**.
Webhooks can be utilized by setting the `callback` object with the URL of your choice, which triggers a `POST` request to the specified webhook URL upon completion of the job process.
```json theme={null}
{
"callback": "https://example.com/callback/url/you/provided"
}
```
### Basic Authentication in Callback URLs
You can specify Basic Authentication credentials directly in the callback URL like this:
```json theme={null}
{
"callback": "https://:@example.com/test/"
}
```
If your password contains special characters (like `@`, `:`, or `/`), be sure to [URL encode](https://developer.mozilla.org/en-US/docs/Glossary/percent-encoding) them.
#### Example (with encoded password)
```json theme={null}
{
"callback": "https://myuser:my%40password@example.com/test/"
}
```
This is useful when the callback endpoint requires HTTP Basic Authentication.
The following APIs accept webhooks:
| Controller | Endpoint |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AI Invoice Parser Controller` | [ai-invoice-parser](/api/ai-invoice-parser) |
| `Barcode Controller` | [barcode/generate](/api/barcode/generate), [barcode/read/from/url](/api/barcode/read) |
| `Convert to PDF Controller` | [xls/convert/to/pdf](/api/convert-from-excel/pdf), [pdf/convert/from/csv](/api/pdf-from-document/csv), [pdf/convert/from/doc](/api/pdf-from-document/doc), [pdf/convert/from/html](/api/pdf-from-html/convert), [pdf/convert/from/image](/api/pdf-from-image), [pdf/convert/from/url](/api/pdf-from-url), [pdf/convert/from/email](/api/pdf-from-email) |
| `Document Parser Controller` | [pdf/documentparser](/api/documentparser/parser) |
| `Email Controller` | [email/send](/api/email/send), [email/extract-attachments](/api/email/extract-attachments), [email/decode](/api/email/decode) |
| `PDF Controller` | [pdf/merge](/api/merge/pdf), [pdf/merge2](/api/merge/various-files), [pdf/split](/api/pdf-split/by-pages), [pdf/split2](/api/pdf-split/by-text-search-or-barcode), [pdf/info](/api/pdf-info-reader), [pdf/info/fields](/api/forms/info-reader), [pdf/find](/api/pdf-find/basic), [pdf/find/table](/api/pdf-find/table), [pdf/security/add](/api/pdf-password/add), [pdf/security/remove](/api/pdf-password/remove), [pdf/classifier](/api/document-classifier), [pdf/attachments/extract](/api/email/extract-attachments) |
| `PDF Data Extraction Controller` | [pdf/convert/to/csv](/api/pdf-to-csv), [pdf/convert/to/html](/api/pdf-to-html), [pdf/convert/to/json2](/api/pdf-to-json/with-ai), [pdf/convert/to/text](/api/pdf-to-text/basic), [pdf/convert/to/text-simple](/api/pdf-to-text/simple), [pdf/convert/to/xls](/api/pdf-to-excel/xls), [pdf/convert/to/xlsx](/api/pdf-to-excel/xlsx), [pdf/convert/to/xml](/api/pdf-to-xml) |
| `PDF Edit Controller` | [pdf/makesearchable](/api/pdf-change-text-searchable/searchable), [pdf/makeunsearchable](/api/pdf-change-text-searchable/unsearchable), [pdf/edit/add](/api/pdf-add), [pdf/edit/rotate](/api/pdf-rotate/basic), [pdf/edit/rotate/auto](/api/pdf-rotate/auto), [pdf/edit/delete-pages](/api/pdf-delete-pages), [pdf/edit/replace-text](/api/pdf-search-text-and-replace/text), [pdf/edit/delete-text](/api/pdf-search-text-and-delete), [pdf/edit/replace-text-with-image](/api/pdf-search-text-and-replace/image) |
| `PDF to Image Controller` | [pdf/convert/to/jpg](/api/pdf-to-image/jpg), [pdf/convert/to/png](/api/pdf-to-image/png), [pdf/convert/to/webp](/api/pdf-to-image/webp), [pdf/convert/to/tiff](/api/pdf-to-image/tiff) |
| `XLS Controller` | [xls/convert/to/csv](/api/convert-from-excel/csv), [xls/convert/to/html](/api/convert-from-excel/html), [xls/convert/to/json](/api/convert-from-excel/json), [xls/convert/to/txt](/api/convert-from-excel/text), [xls/convert/to/xml](/api/convert-from-excel/xml) |
## Summary
Callbacks are functions passed into other functions to be executed after a task is completed, commonly used within the same codebase.
Webhooks are a type of callback that operates over the web, allowing one system to send data to another system when a specific event occurs, typically via an HTTP POST request.
In essence, **webhooks are a practical implementation of the callback concept** across different systems, facilitating real-time communication and integration between web applications.
# Changelog
Source: https://developer.pdf.co/changelog
Notable customer-facing improvements to PDF.co, including new features, API updates, and fixes.
This changelog highlights notable customer-facing improvements to PDF.co. Routine maintenance and internal infrastructure changes are not included.
* Improved AI Invoice Parser job tracking, timeout handling, and result delivery.
* Improved API reliability under high concurrency and repeated job-status requests.
* Added clearer validation and error responses for oversized input downloads.
* Changed API rate-limit windows from per-second to per-minute for smoother request handling.
* Improved AI Invoice Parser result delivery and callback reliability.
* Fixed PDF compression failures involving shared images.
* Improved checkbox flattening and filled-rectangle rendering in PDF processing.
* Improved HTML-to-PDF validation so invalid input and template errors return clearer client errors.
* Improved security protections for HTML-to-PDF conversion.
* Improved HTML-to-PDF support for single-page applications that use URL fragments.
* Improved web-font loading reliability during HTML-to-PDF conversion.
* Added OCR mode selection support to the PDF Split API.
* Improved AI Invoice Parser output consistency for requested custom fields and structured line items.
* Fixed duplicate material numbers caused by adjacent PDF text blocks.
* Improved handling of missing and unrequested custom fields in AI Invoice Parser results.
* Improved AI Invoice Parser extraction for long and multi-page invoices.
* Fixed cases where structured invoice fields could be mixed into unrelated notes.
* Improved timeout and job-cancellation handling for PDF conversion.
* Improved error handling when converting images from remote URLs.
* Added line-item structure hints to AI Invoice Parser.
* Added support for JSON objects in configurable AI Invoice Parser fields.
* Added page counts to AI Invoice Parser responses.
* Improved validation and normalization of AI Invoice Parser line-item structures.
* Improved timeout and aborted-job reporting for asynchronous processing.
* Improved AI Invoice Parser extraction quality, caching, and custom-field handling.
* Improved temporary-file handling and reliability for asynchronous job results.
* Added clearer rate-limit details to API error responses.
* Added more flexible result-retention periods and safer uploaded filename handling.
* Added a legacy rendition option and configurable pre-render delay for HTML and URL conversion.
* Added redaction support to PDF search-and-replace operations.
* Added password support for PDF compression and information extraction.
* Added Data Matrix barcode generation.
* Added support for importing email attachments into PDF documents.
* Improved webhook callback retries and conversion error messages.
* Improved HTML-to-PDF performance and compatibility with a newer browser-rendering engine.
* Improved compatibility with legacy Excel files during Excel-to-PDF conversion.
* Added automatic credit refill support.
* Added PDF Split and PDF Info processing to the newer processing platform.
* Added TIFF and CMYK image support.
* Added Google Drive downloads and presigned file-upload support.
* Added profile support to simple PDF-to-text conversion.
* Improved AI Invoice Parser API validation and result handling.
* Added the v2 PDF compression API.
* Expanded PDF editing with text, images, form fields, checkboxes, radio buttons, and formatting options.
* Added profile support to HTML and URL conversion.
* Added PDF flattening, crop-box handling, and PDF encryption and decryption.
* Added a simple PDF-to-text API with configurable line endings.
* Improved the scalability and reliability of HTML and URL conversion.
* Improved credit refunds for failed and aborted jobs.
* Added batch handling for pay-as-you-go billing.
* Changed webhook processing so callback delivery does not consume additional credits.
* Added Data Store APIs and webhook support for stored data.
* Added the Delay API for scheduled processing workflows.
* Expanded integration-key support across account, file, and job APIs.
* Added monthly credit-pack support.
* Improved optional callbacks and callbacks for aborted jobs.
* Introduced the AI Invoice Parser API and Document Parser tools.
* Added API management for reusable HTML templates.
* Added Google Drive file support to document-processing workflows.
* Added duration and job-duration fields to API responses.
* Added profile support to simple PDF-to-text conversion.
* Added the PDF Inspector as an alias for the PDF editing helper.
* Added rotation-angle support for text watermarks and PDF editing.
* Improved the Document Parser template editor and related APIs.
* Improved PDF-to-image, HTML-to-PDF, and table-detection reliability.
* Added TextSense document-processing integration.
* Added output-link expiration information to API responses.
* Improved OCR and CSV conversion behavior.
* Improved PDF-to-image rendering for logos and object backgrounds.
* Added a general webhook interface for processing callbacks.
* Improved the dashboard, API Logs page, subscription pages, and local-time display.
* Expanded PDF form-field reading and editing for checkboxes, combo boxes, and list boxes.
* Added user-controlled PDF encryption and transparent colors for PDF editing.
* Added direct-download link support for online file-sharing services.
* Added paper-size support to image-to-PDF conversion.
* Improved PDF text replacement, output filenames, and embedded-document handling.
* Added public links to the PDF.co service status page and feature-request portal.
* Added barcode-based PDF splitting.
* Added profile support to PDF optimization.
* Improved right-to-left text search and PDF text extraction.
* Expanded Email Send API options and attachment information.
* Added encrypted file uploads.
* Added the Document Classifier and improved the Document Parser editor.
* Expanded HTML-to-PDF profile controls for input styles.
* Improved PDF information extraction for protected and damaged documents.
# Welcome to PDF.co Docs
Source: https://developer.pdf.co/index
Your all-in-one guide to automate PDF, documents, and data extraction with APIs and integrations.
## Quickstarts
Whether you’re a developer or a no-code builder, you’ll find examples, quickstarts and step-by-step guides to get you up and running fast.
### For API Developers
### For Automation Users
### Basic Knowledge
## Popular Endpoints
Extract structured data from any invoice automatically—no templates needed.
Add text, images, forms, other PDFs, fill forms, links to external sites and external PDF files.
Optimize PDF files up to 13 times smaller in file size by optimizing images and objects.
Extract data from PDFs, JPGs, and PNGs — including fields, tables, values, and barcodes.
# Airtable
Source: https://developer.pdf.co/integrations/airtable
[Airtable](https://airtable.com/) is a cloud\-based project management system that was founded on the belief that software shouldn’t dictate how you work.
## Airtable and PDF.co Plugins
## Airtable and PDF.co Integration via Zapier
## How to Generate PDF Files
# Google Apps Script
Source: https://developer.pdf.co/integrations/google-apps-script
[Google Apps Script](https://developers.google.com/apps-script) is a scripting platform designed for rapid application development for the fast and easy creation of business applications that integrate with **G Suite** products. Modern **JavaScript** is the scripting language being used to write codes. **Apps Script** includes built\-in libraries for **G Suite** applications such as **Drive**, **Calendar**, **Gmail**, and more.
## Apps Script and PDF.co integration
Please contact us to find out more:
## Merge Google Drive PDF Files
## Sample Code
```javascript theme={null}
// Prepare Payload
var data = {
"async": false,
"encrypt": false,
"inline": true,
"name": "result",
"url": pdfUrl
};
// Prepare Request Options
var options = {
'method' : 'post',
'contentType': 'application/json',
'headers': {
"x-api-key": pdfCoAPIKey
},
// Convert the JavaScript object to a JSON string.
'payload' : JSON.stringify(data)
};
// Get Response
// https://developers.google.com/apps-script/reference/url-fetch
var pdfCoResponse = UrlFetchApp.fetch('https://api.pdf.co/v1/pdf/merge', options);
var pdfCoRespContent = pdfCoResponse.getContentText();
var pdfCoRespJson = JSON.parse(pdfCoRespContent);
```
We can break PDF.co API integration into three steps.
* Prepare Payload
* Prepare Request Options.
* Invoke Request and consume the response
It’s worth observing the second step where we’re preparing options for request. Here, **Request Options** contain **API Key** in the header as well payload attribute containing `JSON` string.
The full source code is as below.
```javascript theme={null}
/**
* IMPORTANT: Add Service reference for "Drive". Go to Services > Locate "Drive (drive API)" > Add Reference of it
*/
// Add Your PDF.co API Key here
const pdfCoAPIKey = 'PDFco_API_Key_Here';
// Get the active spreadsheet and the active sheet
ss = SpreadsheetApp.getActiveSpreadsheet();
ssid = ss.getId();
// Look in the same folder the sheet exists in. For example, if this template is in
// My Drive, it will return all of the files in My Drive.
var ssparents = DriveApp.getFileById(ssid).getParents();
// Store File-Ids/PermissionIds used for merging
let filePermissions = [];
/**
* Note: Here, we're getting current folder where spreadsheet is residing.
* But we can certainly pick any folder of our like by using Folder related functions.
* For example:
var allFolders = DriveApp.getFoldersByName("Folder_Containing_PDF_Files");
while (allFolders.hasNext()) {
var folder = allFolders.next();
Logger.log(folder.getName());
}
*/
// Loop through all the files and add the values to the spreadsheet.
var folder = ssparents.next();
/**
* Add PDF.co Menus in Google Spreadsheet
*/
function onOpen() {
var menuItems = [
{name: 'Merge All Files From Current Folder', functionName: 'mergePDFDocumentsFromCurrentFolder'}
];
ss.addMenu('PDF.co', menuItems);
}
function mergePDFDocumentsFromCurrentFolder(){
var allFilesLink = getPDFFilesFromCurFolder(pdfCoAPIKey);
mergePDFDocuments(allFilesLink, pdfCoAPIKey);
}
/**
* Get all PDF files from current folder
*/
function getPDFFilesFromCurFolder(pdfCoAPIKey) {
var files = folder.getFiles();
var allFileUrls = [];
while (files.hasNext()) {
var file = files.next();
var fileName = file.getName();
if(fileName.endsWith(".pdf")){
// Create Pre-Signed URL from PDF.co
var respPresignedUrl = getPDFcoPreSignedURL(fileName, pdfCoAPIKey)
if(!respPresignedUrl.error){
var fileData = file.getBlob();
if(uploadFileToPresignedURL(respPresignedUrl.presignedUrl, fileData, pdfCoAPIKey)){
// Add Url
allFileUrls.push(respPresignedUrl.url);
}
}
}
}
return allFileUrls.join(",");
}
/**
* Merges PDF URLs using PDF.co and Save to drive
*/
function mergePDFDocuments(pdfUrl, pdfCoAPIKey) {
// Get Cells for Input/Output
let resultUrlCell = ss.getRange("A4");
// Prepare Payload
var data = {
"async": false,
"encrypt": false,
"inline": true,
"name": "result",
"url": pdfUrl
};
// Prepare Request Options
var options = {
'method' : 'post',
'contentType': 'application/json',
'headers': {
"x-api-key": pdfCoAPIKey
},
// Convert the JavaScript object to a JSON string.
'payload' : JSON.stringify(data)
};
// Get Response
// https://developers.google.com/apps-script/reference/url-fetch
var pdfCoResponse = UrlFetchApp.fetch('https://api.pdf.co/v1/pdf/merge', options);
var pdfCoRespContent = pdfCoResponse.getContentText();
var pdfCoRespJson = JSON.parse(pdfCoRespContent);
// Display Result
if(!pdfCoRespJson.error){
// Upload file to Google Drive
uploadFile(pdfCoRespJson.url);
// Update Cell with result URL
resultUrlCell.setValue(pdfCoRespJson.url);
}
else{
resultUrlCell.setValue(pdfCoRespJson.message);
}
}
/**
* Gets PDF.co Presigned URL
*/
function getPDFcoPreSignedURL(fileName, pdfCoAPIKey){
// Prepare Request Options
var options = {
'method' : 'GET',
'contentType': 'application/json',
'headers': {
"x-api-key": pdfCoAPIKey
}
};
var apiUrl = `https://api.pdf.co/v1/file/upload/get-presigned-url?name=${fileName}`;
// Get Response
// https://developers.google.com/apps-script/reference/url-fetch
var pdfCoResponse = UrlFetchApp.fetch(apiUrl, options);
var pdfCoRespContent = pdfCoResponse.getContentText();
var pdfCoRespJson = JSON.parse(pdfCoRespContent);
return pdfCoRespJson;
}
/**
* Uploads File to PDF.co PreSigned URL
*/
function uploadFileToPresignedURL(presignedUrl, fileContent, pdfCoAPIKey){
// Prepare Request Options
var options = {
'method' : 'PUT',
'contentType': 'application/octet-stream',
'headers': {
"x-api-key": pdfCoAPIKey
},
// Convert the JavaScript object to a JSON string.
'payload' : fileContent
};
// Get Response
// https://developers.google.com/apps-script/reference/url-fetch
var pdfCoResponse = UrlFetchApp.fetch(presignedUrl, options);
if(pdfCoResponse.getResponseCode() === 200){
return true;
}
else{
return false;
}
}
/**
* Save file URL to specific location
*/
function uploadFile(fileUrl) {
var fileContent = UrlFetchApp.fetch(fileUrl).getBlob();
folder.createFile(fileContent);
}
```
If we quickly analyze the above source code, we’re doing the following.
* Adding necessary references such as “Drive”, as well declaring constant for **PDF.co API Key**. We’re also getting references for the input **Google Drive** folder which contains input files.
* In the Next Step, we’re iterating through all **PDF** files from a folder and uploading them to the **PDF.co** cloud. After the file is uploaded to **PDF.co**, we’re preparing an array of **PDF.co** cloud URLs.
* Then we’re performing a merge using **PDF.co** URLs and saving result files to the output folder.
# PDF.co Integrations
Source: https://developer.pdf.co/integrations/index
PDF.co can be integrated with a lot of online applications, services and platforms with little or no coding experience required.
Automate your tasks and manage your workflows to achieve your goals, for example, automatically trigger events on documents or emails as they arrive in your inbox or easily create PDF invoices from templates.
# Add Password and Security into PDF
Source: https://developer.pdf.co/integrations/make/add-security-to-pdf
This feature enables the addition of password protection and security settings to a **PDF** document.
**Modifying Restriction Settings**: To modify assembly or extraction settings that control PDF restrictions (such as `allowPrintDocument`, `allowFillForms`, `allowModifyDocument`, `allowAssemblyDocument`, and related permission settings), you **must** use the `ownerPassword` in your request. The `userPassword` alone cannot modify these permission settings. Attempting to change these restrictions with only a `userPassword` will result in an error: `"This file is password-protected. Please ensure you've entered the correct password…"`. This requirement applies only when making changes to those restrictions.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Owner Password** | Set the main owner password for document encryption and managing restrictions. | No |
| **User Password** | Optional user password for viewing and printing the document. | No |
| **Encryption Algorithm** | Choose the encryption algorithm. Options: `RC4_40bit`, `RC4_128bit`, `AES_128bit`, `AES_256bit`. Recommended: `AES_128bit` or higher. | No |
| **Print Document** | Choose to allow or prohibit printing the **PDF** document. | No |
| **Print Quality** | Define allowed printing quality. | No |
| **Assembly Document** | Allow or prohibit assembling the document. | No |
| **Content Extraction** | Allow or prohibit copying content from the **PDF** document. | No |
| **Accessibility Support** | Allow or prohibit accessibility features in the **PDF** document. | No |
| **Modify Document** | Allow or prohibit modification of the **PDF** document. | No |
| **Fill Forms** | Allow or prohibit filling of interactive form fields (including signature fields) in the **PDF** document. | No |
| **Modify Annotations** | Allow or prohibit interacting with text annotations and forms in the **PDF** document. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Add Text and Images To a PDF
Source: https://developer.pdf.co/integrations/make/add-text-images-formfields-to-pdf
This feature allows for the addition of text and images to a **PDF** document.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. Leave empty to create a new **PDF** file. | No |
| **Output File Name** | Specify a custom file name for the output file. | No |
## Parameters
### Text Annotations
| Name | Description | Required |
| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Text** | Enter the text for the text object. Macros like line breaks (`\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. For using special macros on **Make.com**, see [Macros for Text](/knowledgebase/macros-for-text#special-macro-style%3A-square-brackets). | No |
| **Pages** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
| **Font Size** | Specify the font size for the text. | No |
| **Font Italic** | Set the text to italic style. | No |
| **Font Bold** | Set the text to bold style. | No |
| **Font Strikeout** | Add a strikeout effect to the text. | No |
| **Font Underline** | Underline the text. | No |
| **Font Name** | Specify the font name. | No |
| **Font Color** | Set the font color using HTML color codes, e.g., `CCBBAA`. | No |
| **Link** | Add an optional clickable link (starting with `http://`, `https://`, `mailto:name@example.com`, etc.). | No |
| **Transparent** | Set the text background as transparent. | No |
| **Width** | Define the width of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Height** | Define the height of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Alignment** | Set text alignment as `Center`, `Right`, or `Left`. Default is `Center`. | No |
| **Type** | Choose the object type: regular text, input control fields, or checkboxes. | No |
| **Id** | Optional. For input fields (text fields or checkboxes), set the field name (or `id`). | No |
### Images
**Images**
| Name | Description | Required |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **URL to the source image.** | Provide a URL to the image, a base64 encoded image, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). | Yes |
| **Pages or pages range.** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Use `!` for reverse indices, e.g. `!0` for the last page and `!1` for the second last page. | No |
| **Link** | Optional link (`http://`, `https://`, `mailto:info@example.com` or similar) to open on click. | No |
| **Width** | Specify the width for the image. Leave empty for automatic detection. | No |
| **Height** | Specify the height for the image. Leave empty for automatic detection. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Template Data** | Use optional **JSON** data to reference inside annotations and fields, for example, `[[variable1]]` with **JSON** data like `{ 'variable1': 'hey hey'}`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | -------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew. |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
You can use the `Custom Profiles` parameters above to:
* [Crop a PDF File](https://developer.pdf.co/api/pdf-add#crop-a-pdf-file)
* [Disable Ligaturization](https://developer.pdf.co/api/pdf-add#disable-ligaturization)
* [Flatten Document](https://developer.pdf.co/api/pdf-add#flatten-document)
* Visit this page for [general information on Profiles usage.](https://developer.pdf.co/api/profiles)
# AI Invoice Parser
Source: https://developer.pdf.co/integrations/make/ai-invoice-parser
Utilize the **PDF.co** AI Invoice Parser to automatically detect invoice and structurally extract data with our advancedu AI. The AI Invoice parser automatically detects invoice layouts without the manual effort previously required to supply document parsing templates for reference.
**Important**
* **Only invoices will be parsed**. For all other documents, please use our existing [**Document Parser**](/integrations/make/document-parser).
* To ensure accurate processing, each invoice must be clearly separated. **If an invoice contains multiple pages, we recommend splitting it** into individual PDFs using the [PDF Split ](/integrations/make/split-pdf)node in your workflow.
* While AI Invoice Parser supports multi-page invoices, **the total page count for a single PDF must not exceed 100 pages**. Submitting large PDFs containing multiple invoices is not recommended.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a file from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
**Import a file from URL**
| Name | Description | Required |
| ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| Name | Description | Required |
| --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **customField** | Comma-separated list of [custom field](/integrations/make/ai-invoice-parser#custom-fields) names to extract. Use `camelCase` for field names (e.g., `storeNumber`, `deliveryDate`). | No |
| **lineItemStructure** | A JSON object that defines a [custom structure](/integrations/make/ai-invoice-parser#line-item-structure) for line items. Each key is a field name (in `camelCase`) and each value is `"string"` or `"number"`. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
### Custom Fields
AI Invoice Parser with custom fields support automatically detects invoice layouts and extracts both standard schema data and user-specified custom fields without requiring manual templates.
The `customField` parameter allows you to specify additional fields to extract beyond the standard schema. Some examples include:
* `storeNumber` - Store or branch identifier
* `deliveryDate` - Expected delivery date
* `financialCharges` - Additional financial charges
* `lineTotal` - Total amount for line items
* `purchaseOrderRef` - Purchase order reference number
* `customerReference` - Customer reference number
* `departmentCode` - Department or cost center code
If a custom field returns an empty value, please [contact our support team](https://pdf.co/support/request?subject=ai-invoice-parser%20-%20custom%20fields) to help improve the extraction accuracy.
### Line Item Structure
The `lineItemStructure` parameter lets you define a custom schema for line items. Each key is a field name you choose (in `camelCase`) and each value is the expected data type — either `"string"` or `"number"`.
When provided, every object inside the `lineItems` array will contain exactly the fields you specified. If a value cannot be extracted from the invoice, the field will still be present with an empty or default value instead of being omitted.
Use `camelCase` for field names (e.g., `unitPrice`, `totalPrice`). The field names you define will be used as-is in the response, giving you full control over the output keys.
## Output
| Name | Description |
| ------------------ | ------------------------------------------------------------------------------ |
| `body` | An object array containing the all invoice parsing result. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
| `Page Count` | Total Page Count. |
## Supported Languages
* **Albanian (Shqip)**
* **Bosnian (Bosanski)**
* **Bulgarian (Български)**
* **Croatian (Hrvatski)**
* **Czech (Čeština)**
* **Danish (Dansk)**
* **Dutch (Nederlands)**
* **English**
* **Estonian (Eesti)**
* **Finnish (Suomi)**
* **French (Français)**
* **German (Deutsch)**
* **Greek (Ελληνικά)**
* **Hungarian (Magyar)**
* **Icelandic (Íslenska)**
* **Italian (Italiano)**
* **Latvian (Latviešu)**
* **Lithuanian (Lietuvių)**
* **Norwegian (Norsk)**
* **Polish (Polski)**
* **Portuguese (Português)**
* **Romanian (Română)**
* **Russian (Русский)**
* **Serbian (Српски)**
* **Slovak (Slovenčina)**
* **Slovenian (Slovenščina)**
* **Spanish (Español)**
* **Swedish (Svenska)**
* **Turkish (Türkçe)**
* **Ukrainian (Українська)**
# Generate a Barcode
Source: https://developer.pdf.co/integrations/make/barcode-generate
This functionality allows for the generation of high\-quality, printable, and scannable barcodes in image or **PDF** formats. It supports a wide range of barcode types including **QR Code**, **Code 39**, **Code 128**, **Datamatrix**, **PDF417**, **UPC**, **EAN**, and many others.
## Input
| Name | Description | Required |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Barcode Value** | Specify the value to be encoded into the barcode. | Yes |
| **Barcode Type** | Choose the type of barcode to generate. Defaults to **QR Code**, with various other formats available. | No |
| **Output File Name** | Specify a custom file name for the output file. | No |
| **Inline** | Set to `true` to generate URL as inline `datauri` link that you can embed directly into HTML. Important: you need to switch to `JSON` as output instead of `Download File` so it will generate a link with datauri inline data. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
***
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `Angle` | integer | `0` | See [profiles.Angle](/api/barcode/generate#profiles-angle) |
| `NarrowBarWidth` | integer | `3` | See [profiles.NarrowBarWidth](/api/barcode/generate#profiles-narrowbarwidth) |
| `CaptionFont` | string | `Arial, 12` | See [profiles.CaptionFont](/api/barcode/generate#profiles-captionfont) |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Read a Barcode
Source: https://developer.pdf.co/integrations/make/barcode-read
This feature facilitates the reading of barcodes from various sources such as images, **TIFF** files, **PDF** documents, and scanned documents. It is capable of interpreting all popular barcode types, ranging from **Code 39** and **Code 128** to **QR Code**, **Datamatrix**, and **PDF417**. The tool is designed to effectively handle noisy and damaged barcodes, scans, and documents, ensuring accurate barcode recognition.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF or image from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF or image from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Barcode Type** | Select the barcode type for decoding. Defaults to **QR Code**, with support for various other formats. | No |
| **Pages** | Specify page numbers or ranges for barcode reading. Leave blank to scan all pages. The first page starts at `0`. Example: `0,2-5,7-`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Barcodes` | An array containing detailed barcode information such as `Value`, `Type`, `TypeName`, `Page`, `Rect`, and others. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Compress and Optimize PDF
Source: https://developer.pdf.co/integrations/make/compress-pdf
This feature is designed to optimize and compress **PDF** documents effectively. It reduces the file size of your **PDF** while maintaining high\-quality standards. This optimization is essential for managing large **PDF** files, ensuring they are easier to share and use without compromising on visual quality.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------ | ------- | ------- | ------------------------------------------------------------------------------------------------------------ |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `JPEGQuality` | integer | 25 | Controls JPEG compression quality from 1 (worst quality, smallest size) to 100 (best quality, largest size). |
# Convert from PDF
Source: https://developer.pdf.co/integrations/make/convert-from-pdf
This feature offers versatile functionality for converting **PDF** pages into various formats. It supports conversion to structured **CSV**, **XML**, **JSON**, plain text, and image formats such as **JPG**, **PNG**, and **TIFF**, catering to a wide range of needs for **PDF** content manipulation.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Convert Type** | Select the type of conversion. Default is `PDF to Text`. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave empty to include all pages. The first page is numbered `0` (zero). For example: `0,1-2,5-`. | No |
| **Password** | If the **PDF** is password-protected, enter the password here. | No |
| **Inline** | Choose `Yes` to include a copy of the output data in `JSON output` mode. If set to `No`, only a link to the output will be included in this mode. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Body` | Represents the raw output data. This is generated only when the `Export Type` option is set to `JSON Output`. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description | Available for |
| ------------------------------ | ----------------------------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64`. | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `ColumnDetectionMode` | string | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See Column Detection Mode for more information. | PDF to CSV, PDF to XLS |
| `OCRMode` | string | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OCRResolution` | integer | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `LineGroupingMode` | string | None | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `ConsiderFontColors` | boolean | false | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DetectNewColumnBySpacesRatio` | string | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `AutoAlignColumnsToHeader` | boolean | true | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RenderTextObjects` | boolean | true | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See Disable Text Layer for more information. | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RenderImageObjects` | boolean | true | Render image objects or not. | PDF to JPG, PDF to PNG, PDF to WEBP |
| `RenderVectorObjects` | boolean | true | Render vector objects or not. | PDF to JPG, PDF to PNG, PDF to WEBP |
| `JPEGQuality` | integer | 85 | See profiles.JPEGQuality. | PDF to JPG |
| `WEBPQuality` | integer | 75 | See profiles.WEBPQuality. | PDF to WEBP |
| `TIFFCompression` | string | LZW | See profiles.TIFFCompression. | PDF to TIFF |
| `RenderingResolution` | integer | 120 | See Set Image Resolution for more information. | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OptimizeImages` | boolean | true | Some PDF may have high quality images used in the document and you may need to keep the quality of these images in the output HTML. By default PDF to HTML is optimizing images and you can easily turn it off. See Control Image Quality for more information. | PDF to HTML |
| `OutputPageWidth` | integer | 1024 | Control page width (in pixels) for output HTML. Height is calculated and used according to the original pdf pages ratio. See Control Output Page Width for more information. | PDF to HTML |
| `AdditionalCssStyles` | string | "" | To inject CSS for layout options in your HTML. Example: `#canvas { zoom: 50%; }`. Scale the div that contains all generated HTML pages by 50%. See Inject CSS for more information. | PDF to HTML |
| `SaveVectors` | boolean | false | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `SaveImages` | string | None | Controls how images are saved during PDF to HTML conversion. Modes: None (no images), OuterFile (save to sub-folder), Embed (embed as Base64 data:URI). | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML, PDF to HTML |
| `ConsiderFontSizes` | boolean | false | Set to true to make the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `ExtractionArea` | array\[number] | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `ExtractShadowLikeText` | boolean | true | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to Auto to properly apply the shadow text filtering effect. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
# Convert into PDF
Source: https://developer.pdf.co/integrations/make/convert-to-pdf
This feature allows for the conversion of a variety of file types into **PDF** format. It supports transforming documents, spreadsheets, presentations, emails, and single images into high\-quality **PDF** files, making it versatile for different conversion needs.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Convert Type** | Select the type of conversion. Default is `Document to PDF`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Create Fillable PDF Form
Source: https://developer.pdf.co/integrations/make/create-fillable-pdf-form
This feature allows for the addition of form fields to a **PDF** document.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. Leave empty to create a new **PDF** file. | No |
| **Output File Name** | Specify a custom file name for the output file. | No |
## Parameters
### Form Field
| Name | Description | Required |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Type** | Choose the object type: text input control, multiline text input, or checkboxes. | Yes |
| **Id** | Optional. For input fields (text fields or checkboxes), set the field name (or `id`). | Yes |
| **Width** | Define the width of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Height** | Define the height of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Text** | Enter the text for the text object. Macros like line breaks (`\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. | No |
| **Pages** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
### Text
| Name | Description | Required |
| ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Text** | Enter the text for the text object. Macros like line breaks (`\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. | No |
| **Pages** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
| **Font Size** | Specify the font size for the text. | No |
| **Font Italic** | Set the text to italic style. | No |
| **Font Bold** | Set the text to bold style. | No |
| **Font Strikeout** | Add a strikeout effect to the text. | No |
| **Font Underline** | Underline the text. | No |
| **Font Name** | Specify the font name. | No |
| **Font Color** | Set the font color using HTML color codes, e.g., `CCBBAA`. | No |
| **Link** | Add an optional clickable link (starting with `http://`, `https://`, `mailto:name@example.com`, etc.). | No |
| **Transparent** | Set the text background as transparent. | No |
| **Width** | Define the width of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Height** | Define the height of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Alignment** | Set text alignment as `Center`, `Right`, or `Left`. Default is `Center`. | No |
### Images
| Name | Description | Required |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **URL to the source image.** | Provide a URL to the image, a base64 encoded image, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). | Yes |
| **Pages or pages range.** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Use `!` for reverse indices, e.g. `!0` for the last page and `!1` for the second last page. | No |
| **Link** | Optional link (`http://`, `https://`, `mailto:info@example.com` or similar) to open on click. | No |
| **Width** | Specify the width for the image. Leave empty for automatic detection. | No |
| **Height** | Specify the height for the image. Leave empty for automatic detection. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Template Data** | Use optional **JSON** data to reference inside annotations and fields, for example, `[[variable1]]` with **JSON** data like `{ 'variable1': 'hey hey'}`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](/api/profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
# Document Classifier
Source: https://developer.pdf.co/integrations/make/document-classifier
Leverage AI to analyze the text of input documents and classify them into categories such as invoices, orders, or industry\-specific types. This feature is particularly useful for quickly identifying the source of a document. It also allows for the implementation of custom classification rules to tailor the analysis to specific needs.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF or Image from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
**Import PDF or Image from URL**
| Name | Description | Required |
| ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| Name | Description | Required |
| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Set custom rules** | Optionally, define classification rules in **CSV** format. Each row should be formatted as `classname,logic,keyword1,keyword2`. Example: `Amazon,AND,Amazon AWS,AWS Invoice`. For detailed instructions, refer to [PDF Classifier](https://pdf.co/pdf-classifier). | No |
| **Load custom rules from CSV via url** | Provide a link to a **CSV** containing custom classification rules. Each row should be formatted as `classname,logic,keyword1,keyword2`. Example: `Amazon,AND,Amazon AWS,AWS Invoice`. For detailed instructions, refer to [PDF Classifier](https://pdf.co/pdf-classifier). | No |
| **Case Sensitive Custom Rules Enabled** | Specify whether the keywords in custom rules should be case sensitive. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Body` | Contains the identified document categories, listed in a `classes` string array. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------ | ----------------------------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RenderTextObjects` | boolean | `true` | Render text objects or not |
| `RenderVectorObjects` | boolean | `true` | Render vector objects or not |
| `RenderImageObjects` | boolean | `true` | Render image objects or not |
| `TIFFCompression` | string | `LZW` | TIFF compression algorithm. The options are: None, LZW, CCITT3, CCITT4, RLE |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | `["1.4"]` | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
You can also use `Custom Profiles` to:
* [Disable text layer](https://developer.pdf.co/api/pdf-to-image/png#disable-text-layer).
* [Set image preprocessing filters.](https://developer.pdf.co/api/pdf-to-image/png#ocrimagepreprocessingfilters)
* Visit this page for [general information on Profiles usage.](https://developer.pdf.co/api/profiles)
# Parse a Document
Source: https://developer.pdf.co/integrations/make/document-parser
Utilize the **PDF.co** Document Parser to automatically extract information from various documents such as invoices, reports, orders, and statements. This feature supports both built\-in and custom extraction templates, facilitating efficient and accurate data retrieval from fields and tables in documents.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a file from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
**Import a file from URL**
| Name | Description | Required |
| ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| Name | Description | Required |
| ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Document Parser Template ID** | Use `1` for the built-in Invoice Parser template or specify custom template IDs. Manage your Document Parser templates at [Document Parser Template Editor](https://app.pdf.co/document-parser/templates). | No |
| **Custom Template Code** | For on-premise installations, enter the Custom Document Parser Template Code. | No |
| **Output Format** | Choose `JSON` for JSON output, `CSV` for comma-separated values, or `XML`. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave blank for all pages. The first page is `0` (zero). For example: `0,1-2,5-`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Body` | Delivers a parsed object array with results formatted as `Name`, `Value`, and `Object Type`. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Fill a PDF Form
Source: https://developer.pdf.co/integrations/make/fill-pdf-form
This feature is designed to fill out **PDF** forms automatically.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. Leave empty to create a new **PDF** file. | No |
| **Output File Name** | Specify a custom file name for the output file. | No |
## Parameters
### Text Annotations
| Name | Description | Required |
| ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Text** | Enter the text for the text object. Macros like line breaks (`\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. | No |
| **Pages** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
| **Font Size** | Specify the font size for the text. | No |
| **Font Italic** | Set the text to italic style. | No |
| **Font Bold** | Set the text to bold style. | No |
| **Font Strikeout** | Add a strikeout effect to the text. | No |
| **Font Underline** | Underline the text. | No |
| **Font Name** | Specify the font name. | No |
| **Font Color** | Set the font color using HTML color codes, e.g., `CCBBAA`. | No |
| **Link** | Add an optional clickable link (starting with `http://`, `https://`, `mailto:name@example.com`, etc.). | No |
| **Transparent** | Set the text background as transparent. | No |
| **Width** | Define the width of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Height** | Define the height of the text box. Coordinates start at the top left (use the provided viewer to measure coordinates). | No |
| **Alignment** | Set text alignment as `Center`, `Right`, or `Left`. Default is `Center`. | No |
| **Type** | Choose the object type: regular text, input control fields, or checkboxes. | No |
| **Id** | Optional. For input fields (text fields or checkboxes), set the field name (or `id`). | No |
### Images
**Images**
| Name | Description | Required |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **X** | Determine the X coordinate for text placement. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **Y** | Specify the Y coordinate. Use **PDF.co** [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure **PDF** coordinates. | Yes |
| **URL to the source image.** | Provide a URL to the image, a base64 encoded image, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). | Yes |
| **Pages or pages range.** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
| **Link** | Optional link (`http://`, `https://`, `mailto:info@example.com` or similar) to open on click. | No |
| **Width** | Specify the width for the image. Leave empty for automatic detection. | No |
| **Height** | Specify the height for the image. Leave empty for automatic detection. | No |
### Fields
**Fields**
| Name | Description | Required |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Field Name** | Specify the exact names of the form fields that you intend to fill. | Yes |
| **Text** | Enter the text for the text object. Macros like line breaks (`\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. | No |
| **Page or pages range.** | Default is `0` (first page). Use comma-separated ranges for multiple pages like `0,1-2,5,7-`. `7-` means from the 7th to the last page. Negative pages like `-2` for second last page. | No |
| **Font Size** | Specify the font size for the text. | No |
| **Font Italic** | Set the text to italic style. | No |
| **Font Bold** | Set the text to bold style. | No |
| **Font Strikeout** | Add a strikeout effect to the text. | No |
| **Font Underline** | Underline the text. | No |
| **Font Name** | Specify the font name. | No |
| **Font Color** | Set the font color using HTML color codes, e.g., `CCBBAA`. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Template Data** | Use optional **JSON** data to reference inside annotations and fields, for example, `[[variable1]]` with **JSON** data like `{ 'variable1': 'hey hey'}`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | -------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew, use the following: |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Getting Started with Make
Source: https://developer.pdf.co/integrations/make/getting-started
[Make](https://www.make.com) is a visual platform that lets you design, build, and automate anything – from simple tasks to complex workflows – in minutes. With **Make**, you can send information between **PDF.co** and thousands of apps to improve document workflows and data extraction processes seamlessly. It’s fast and easy to use, visually intuitive and requires zero coding expertise.
## How To Connect PDF.co to Make
* Log in to your [PDF.co](https://pdf.co) account. If new to **PDF.co**, [Sign Up here](https://app.pdf.co/signup).
* Click Your API Key and copy the API key to your clipboard.
* Go to **Make** and open the **PDF.co** module's `Create a connection` dialog.
* In the Connection name field, enter a name for the connection.
* The connection has been established.
### Convert DOC to PDF using Make and PDF.co
# Convert HTML to PDF
Source: https://developer.pdf.co/integrations/make/html-to-pdf
This feature allows for the conversion of **HTML** code, **HTML** templates, or entire web pages (URLs) into **PDF** format. It’s designed to transform web content into a portable and universally accessible **PDF** format with ease.
This feature allows for the conversion of **HTML** code, **HTML** templates, or entire web pages (URLs) into **PDF** format. It's designed to transform web content into a portable and universally accessible **PDF** format with ease.
## Input
| Name | Description | Required |
| ------------------- | ------------------------------------------------------------------- | -------- |
| **Convert Options** | Choose from `HTML to PDF`, `HTML Template to PDF`, or `URL to PDF`. | Yes |
***
**HTML to PDF**
| Name | Description | Required |
| ------------------- | ----------------------------------------- | -------- |
| **Input HTML Code** | Provide the raw HTML code for conversion. | Yes |
**HTML Template to PDF**
| Name | Description | Required |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **HTML Template ID** | Enter the ID from [HTML to PDF Templates](https://app.pdf.co/templates/html). | No |
| **Input data in CSV or JSON format** | Supply data in `CSV` or `JSON` format for the selected HTML template. Examples: CSV: `invoice_id, totalrn12345,$999` JSON: `{"invoice_id":"12345","total":"$999"}`. Test your [HTML to PDF Templates](https://app.pdf.co/templates/html). | Yes |
**URL to PDF**
| Name | Description | Required |
| ------------------------- | ---------------------------------------------------------------------------- | -------- |
| **Input URL to web page** | Enter the URL of the web page for conversion. Example: `https://google.com`. | Yes |
| Name | Description | Required |
| ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Output File Name** | Specify a custom file name for the output file. | No |
| **Orientation** | Select the **PDF** orientation: `Portrait` or `Landscape`. | No |
| **Paper Size** | Choose the paper size, such as `Legal`, `Letter`, `A0`, `A1`, `A2`, `A3`, `A4`, `A5`, etc. | No |
| **Render Page Background** | Specify if the page background should be rendered in the **PDF**. | No |
| **Do not wait for full load** | Select if the conversion should proceed without waiting for the full page load. | No |
| **Margins** | Set custom margins, overriding CSS margin styles for the converted page. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `HTMLCodeHeadInject` | string | - | Injects CSS into the HTML `` section to prevent page breaks within specified elements. See HTML to PDF Knowledge Base for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Convert from Images into PDF
Source: https://developer.pdf.co/integrations/make/images-to-pdf
This feature enables the creation of **PDF** documents from one or more image files. It supports **JPEG**, **PNG**, and **TIFF** formats, offering a convenient way to consolidate images into a single **PDF** file.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------- | -------- |
| **Import Options** | Select the input source: `Upload File(s)` or `Input Link(s)`. | Yes |
***
**Upload File(s)**
| Name | Description | Required |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **File Name** | Specify a custom file name for the output file. | No |
**Input Link(s)**
| Name | Description | Required |
| -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input Link** | Enter URLs to source images (e.g., `example1.com/file1.png,example2.com/file2.jpg`), or use a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like **Google Drive** or **Dropbox**, ensure the link is publicly accessible. | Yes |
| **File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description | Available for |
| ------------------------------ | ----------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 | PDF to JPG, PDF to PNG |
| `OCRMode` | string | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. | PDF to JPG, PDF to PNG |
| `OCRResolution` | integer | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. | PDF to JPG, PDF to PNG |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. | PDF to JPG, PDF to PNG |
| `LineGroupingMode` | string | None | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). | PDF to JPG, PDF to PNG |
| `ConsiderFontColors` | boolean | false | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | PDF to JPG, PDF to PNG |
| `DetectNewColumnBySpacesRatio` | string | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | PDF to JPG, PDF to PNG |
| `AutoAlignColumnsToHeader` | boolean | true | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | PDF to JPG, PDF to PNG |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | PDF to JPG, PDF to PNG |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. | PDF to JPG, PDF to PNG |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. | PDF to JPG, PDF to PNG |
| `RenderTextObjects` | boolean | true | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See Disable Text Layer for more information. | PDF to JPG, PDF to PNG |
| `RenderImageObjects` | boolean | true | Render image objects or not | PDF to JPG, PDF to PNG |
| `RenderVectorObjects` | boolean | true | Render vector objects or not | PDF to JPG, PDF to PNG |
| `JPEGQuality` | integer | 85 | See profiles.JPEGQuality | PDF to JPG |
| `RenderingResolution` | integer | 120 | See Set Image Resolution for more information. | PDF to JPG, PDF to PNG |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to JPG, PDF to PNG |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. | PDF to JPG, PDF to PNG |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. | PDF to JPG, PDF to PNG |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to JPG, PDF to PNG |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. | PDF to JPG, PDF to PNG |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. | PDF to JPG, PDF to PNG |
# Integrating File Sources with PDF.co
Source: https://developer.pdf.co/integrations/make/input-file-sources
Supercharge your **Make** workflows by seamlessly integrating with leading file storage services like **Google Drive**, **Dropbox**, **OneDrive**, and **Box**. This guide meticulously outlines how to use these services in harmony with **PDF.co**’s **Make** plugin, enhancing your document management capabilities.
Supercharge your **Make** workflows by seamlessly integrating with leading file storage services like **Google Drive**, **Dropbox**, **OneDrive**, and **Box**. This guide meticulously outlines how to use these services in harmony with **PDF.co**'s **Make** plugin, enhancing your document management capabilities.
## URL Accessibility for PDF.co
The API supports TLS 1.2 and 1.3 for secure connections. Earlier versions such as TLS 1.0 and 1.1 are deprecated and should be avoided.
For successful integration, any publicly accessible URL can be used. Whether you're using services like Google Drive, Dropbox, or others, ensure the URLs are accessible to PDF.co. We recommend using **Make.com**-generated URLs (detailed in the sections below) as they are guaranteed to be compatible with PDF.co. If you're using direct URLs not provided via Make.com, please verify their public accessibility to ensure PDF.co can process your files efficiently.
## Google Drive
Effortlessly integrate **Google Drive** into your workflows:
\- Set up a **Google Drive** Trigger to `Watch Files in a Folder`
\- Follow with **Google Drive**'s `Download a File` action to fetch file binary data for subsequent steps.
\- Use the acquired data as input for **PDF.co** actions.
Here's the complete scenario with all steps:
In the `Download a File` step, utilize the `File ID` from the Trigger as shown below:
Finally, feed the data from `Download a File` into your **PDF.co** action:
## Dropbox
Integrate **Dropbox** with ease:
* Start with the **Dropbox** Trigger to `Watch Files`.
* Add **Dropbox**'s `Download a File` action for acquiring file data.
* Direct this data into **PDF.co**'s action steps.
Visualize the entire flow:
When setting up `Download a File`, use `Path lower` from the Trigger:
The data from `Download a File` now serves as input for **PDF.co** actions:
## OneDrive
Integrating **OneDrive** is straightforward:
* Use the **OneDrive** Trigger to `Watch Files/Folders`.
* Proceed with **OneDrive**'s `Download a File` action.
* Channel the downloaded data into **PDF.co**'s functionalities.
Here's the setup in its entirety:
Configure `Download a File` using `Item ID` from the Trigger:
Finally, use the `Download a File` data as input for **PDF.co**:
## Box
Follow these steps for **Box** integration:
* Set the **Box** Trigger to `Watch Files`.
* Use **Box**'s `Download a File` action to obtain file data.
* This data is then ready for use with **PDF.co** actions.
View the full integration process:
In `Download a File`, fill in the `File ID` using the Trigger's output:
The data from `Download a File` is now primed for **PDF.co** actions:
***
# Job Check
Source: https://developer.pdf.co/integrations/make/job-check
The Job Check feature is designed to check the status of a job initiated on **PDF.co**. This is particularly useful when working with large documents in the `Async For Large Docs` mode, where the processing is done asynchronously.
## Input
| Name | Description | Required |
| ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **PDF.co Job ID** | Input the Job ID as provided by **PDF.co** after initiating a process. | Yes |
| **Auto-Expand JSON, CSV, XML, Text Links to Content** | Enable this option to automatically extract content from the response links (useful for **JSON**, **CSV**, **XML**, Text formats) and integrate it directly into the output. | Yes |
| **Export Type** | Choose between `Download a File` or `Inline`. Default is `Inline`. | No |
***
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Body` | Represents the raw output data. This is generated only when the `Export Type` option is set to `JSON Output`. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
# Make Webhooks Integration with PDF.co
Source: https://developer.pdf.co/integrations/make/make-webhooks
**Make.com** webhooks provide a powerful way to receive real\-time notifications from the **PDF.co** API when specific operations are completed. This integration can significantly enhance your automation workflows in **Make.com** by eliminating the need for polling job statuses via Background Job Check.
**Make.com** [webhooks](https://www.make.com/en/help/tools/webhooks) provide a powerful way to receive real-time notifications from the **PDF.co** API when specific operations are completed. This integration can significantly enhance your automation workflows in **Make.com** by eliminating the need for polling job statuses via [Background Job Check](/api/job-check).
In this guide, we will walk through the process of setting up **Make.com** [webhooks](https://www.make.com/en/help/tools/webhooks) with **PDF.co**, including how to use the [callback](/api/webhooks) parameter via [std\_params](/api/profiles#standard-parameters) in the [profiles](/api/profiles) parameter.
***
## Understanding PDF.co Webhooks
**PDF.co** supports [webhooks](/api/webhooks) for a variety of operations, allowing you to receive notifications when tasks such as document conversions or data extractions are finished. Instead of repeatedly checking the job status, you can set a [callback](/api/webhooks) URL to which **PDF.co** will send a `POST` request with the job results.
To set up a [webhook](/api/webhooks), you can use the [callback](/api/webhooks) parameter directly in your API calls. However, when integrating with automation platforms like **Make.com**, it’s often more convenient to use the [std\_params](/api/profiles#standard-parameters) within the [profiles](/api/profiles) parameter. The [std\_params](/api/profiles#standard-parameters) allows you to define standard API parameters, including the [callback](/api/webhooks) URL, in a JSON format.
For example, a request to **PDF.co** might include:
## Setting Up Make.com Scenarios
If you're using a custom webhook or callback endpoint (rather than Make.com webhooks), it must respond immediately within a few seconds. If it takes longer than **20–30 seconds**, the PDF.co API will mark the callback as failed and retry up to **3 times**.
To integrate **PDF.co** [webhooks](/api/webhooks) with **Make.com**, you’ll need to create two scenarios:
1. **Webhook Scenario**: This scenario receives the [callback](/api/webhooks) from **PDF.co**.
2. **Request Scenario**: This scenario initiates the **PDF.co** operation and sets the callback URL.
## Creating the Webhook Scenario
1. Create a new scenario.
2. Select a [Webhooks](https://www.make.com/en/help/tools/webhooks) module and select [Custom Webhook](https://www.make.com/en/help/tools/webhooks#creating-custom-webhooks) as the trigger.
3. Click “Add” and [configure](https://www.make.com/en/help/tools/webhooks#webhook-settings) to accept `POST` requests, as **PDF.co** API sends callbacks via `POST`.
4. **Make.com** will generate a unique URL for your webhook. Copy this URL for use in the next scenario.
## Initiating the PDF.co Request
1. Create another scenario in **Make.com** with an appropriate trigger (e.g., manual or scheduled).
2. Add a **PDF.co** action module, selecting the desired operation (e.g., Convert From **PDF**).
3. In the action module, click on the “Advance Options” and set the [profiles](/api/profiles) parameter to include the [std\_params](/api/profiles#standard-parameters) with the [callback](/api/webhooks) URL. The [profiles](/api/profiles) parameter should be a string containing a JSON object, like this:
4. Configure other parameters as needed for your specific **PDF.co** operation.
## Processing the Callback
1. In the [webhooks](https://www.make.com/en/help/tools/webhooks) scenario, after the “Custom Webhook” trigger, add actions to process the data received from **PDF.co**.
2. The [callback](/api/webhooks) from **PDF.co** includes details such as the job ID, status, and output URL. Parse this data and integrate it with other apps in **Make.com** as required.
# Merge a PDF
Source: https://developer.pdf.co/integrations/make/merge
This feature allows merging multiple **PDF** files into a single **PDF** document. It also supports a range of other file formats such as **doc**, **docx**, **rtf**, **txt**, **xls**, **xlsx**, **csv**, **jpg**, **png**, and **zip** in Advanced mode.
The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed.
## Input
| Name | Description | Required |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------ | -------- |
| **Enable non-PDF files as input** | Allows **doc**, **docx**, **rtf**, **xls**, **xlsx**, **txt**, **jpg**, **png**, and **zip** files as input. | No |
| **Import Options** | Choose the input source, either `Upload Files` or `Input Links`. | Yes |
***
**Upload Files**
| Name | Description | Required |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **File Name** | Specify a custom file name for the output file. | No |
**Input Links**
| Name | Description | Required |
| -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input Link** | Enter URLs to source files seperated by comma (e.g., `example1.com/file1.pdf,example2.com/file2.pdf`), or use a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like **Google Drive** or **Dropbox**, ensure the link is publicly accessible. | Yes |
| **File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Merging Files from Google Drive using PDF.co
To overcome issues such as restricted files or temporary blocking by Google Drive, use the `Upload a File` step in PDF.co before merging:
1. Set up a **Google Drive** Trigger to `Watch Files in a Folder`.
2. Use **Google Drive**’s `Download a File` action to fetch file binary data.
3. Add the `Upload a File` step in PDF.co to upload the file.
4. Use the uploaded file URL in the `Merge PDF` step.
Here’s the complete scenario with all steps:
In the `Upload a File` step, use the file binary data from the `Download a File` action as shown below:
This approach ensures that the files are accessible and avoids issues related to file restrictions, rate limiting or temporary blocking by Google Drive.
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| --------------------------------- | -------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `RenameMatchingFieldsDuringMerge` | boolean | `true` | This feature enables the renaming of field names during the merging of PDF files which contain forms. If set to false, it will retain the original field names. This is helpful for merged PDF forms with identical field names when the customer wants to auto-fill the identical field names in other pages. |
| `GenerateBookmarks` | boolean | `false` | This adds bookmarks to the merged document with names assigned to every merged document in the same order: |
| `BookmarkTitles` | array\[string] | - | An array containing the titles/names for bookmarks to be created |
| `zipIncludeFilter` | string | - | You can control which files to include and exclude from input zip files with a profiles. |
| `zipExcludeFilter` | string | - | zipIncludeFilter and zipExcludeFilter support `*` and `?` wildcards. |
| `MergedDocumentTitle` | string | Title of the first document | You can change the document title during a merge with the following: |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Get PDF Information
Source: https://developer.pdf.co/integrations/make/pdf-info
This feature retrieves comprehensive information from a **PDF** document. It provides essential details such as the number of pages, author, keywords, and other relevant metadata. This functionality is especially useful for analyzing and cataloging **PDF** documents based on their content and attributes.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
**Import a File From URL**
| Name | Description | Required |
| ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Body` | Contains **PDF** information objects, including `PageCount`, `Title`, `Author`, `Subject`, `CreationDate`, `Creator`, and others. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------ | ----------------------------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | `["1.4"]` | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Convert from PDF into Images
Source: https://developer.pdf.co/integrations/make/pdf-to-images
This feature facilitates the conversion of **PDF** documents into image formats. It allows for the transformation of each selected **PDF** page or page range into either **JPG** or **PNG** format, providing flexibility for various image conversion needs.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Convert Options** | Choose the conversion format. Valid options are `PDF to JPG` and `PDF to PNG`. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for conversion. Leave blank for all pages. The first page is numbered `0` (zero). For example: `0,1-2,5-`. | No |
| **Inline** | Set to `true` to receive the extracted data as a `body` string. Caution: Do not enable this option if you need to download the output file. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](/api/profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `URLs` | Contains string array of image URLs. This data is generated only when the `Export Type` option is set to `JSON Output`. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
# Make PDF.co API Call
Source: https://developer.pdf.co/integrations/make/pdfco-api-call
This feature enables users to perform arbitrary authorized API calls to **PDF.co**. It provides the flexibility to use various **PDF.co** API endpoints for a range of **PDF** related tasks.
## Input
| Name | Description | Required |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **API Endpoint Path** | Specify the path to a **PDF.co** API endpoint, relative to `https://api.pdf.co/v1`. Example: **/pdf/edit/add**. For parameter details, refer to [PDF.co API documentation](/api). Check recent API logs at [API Logs page](https://app.pdf.co/account/logs/api). | Yes |
| **Method** | Choose the HTTP method: `GET`, `POST`, `PUT`, `PATCH`, `DELETE`. | Yes |
| **Headers** | Add additional header Key/Value to the request. Authorization headers are automatically included. | Yes |
***
| Name | Description | Required |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input Type** | Override URL parameter by selecting an Input Type. Options: `None (uses Body only)`, `Upload files and inject as URL param`, `Override URL param with value`. | Yes |
**Upload files and inject as url param**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Override url param with value**
| Name | Description | Required |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| ----------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Body** | Add request body key/value. Refer to [PDF.co API documentation](/api) for details on supported parameters. | No |
| **Forcely enable \`async\` mode** | Forces async mode by adding `async: true` to input params. Recommended for better performance. No effect if `Body` already has `async` defined. | No |
| **Auto-run \`job/check\` for async jobs** | Automatically runs `v1/job/check` for async jobs until completion. Note: Checking async job status consumes credits. For long jobs (40+ sec), disable this and use a separate module for job status checks. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
# Remove Password and Security from PDF
Source: https://developer.pdf.co/integrations/make/remove-security-from-pdf
This function is designed to remove password protection and security features from a **PDF** document. It is useful for unlocking **PDFs** and ensuring unrestricted access to their content.
## Input
| Name | Description | Required |
| ------------------ | ------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import PDF from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import PDF from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Password** | Enter the owner/user password of the **PDF** file to remove its security features. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Search and Delete Found Text in PDF
Source: https://developer.pdf.co/integrations/make/search-and-delete-text
This function allows for the removal of specified text within a **PDF** document. It’s particularly useful for erasing sensitive or unwanted information from **PDF** files.
This function allows for the removal of specified text within a **PDF** document. It's particularly useful for erasing sensitive or unwanted information from **PDF** files.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Text to Search and Delete** | Specify the text that needs to be searched for and deleted in the **PDF**. | No |
| **Use Regular Expressions** | Opt to use regular expressions for more complex search patterns. For instance, to find a **SSN** format, use `[0-9]{3}-[0-9]{2}-[0-9]{4}`. | No |
| **Case Sensitive?** | Determine if the search should be case sensitive. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave empty for all pages. The first page is numbered `0` (zero). Example: `0,1-2,5-`. | No |
| **Password** | If the **PDF** is password-protected, enter the password here. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `JSON Output`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `removeTextUnderPatch` | boolean | `true` | Controls whether to remove text under the patch or not |
| `usepatch` | boolean | `false` | Controls whether to use a patch or not |
| `patchColor` | string | `#000000` | Controls the color of the patch |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Search and Replace Text in PDF
Source: https://developer.pdf.co/integrations/make/search-and-replace-text
This feature allows you to search for specific text within a **PDF** document and replace it with new text. It is particularly useful for updating documents without the need to manually edit the **PDF**.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Search and Replace String**
| Name | Description | Required |
| ------------------------ | ------------------------------------------------------- | -------- |
| **Text to Search** | Specify the text you want to search for in the **PDF**. | Yes |
| **Text to Replace With** | Enter the new text that will replace the found text. | Yes |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Use Regular Expressions** | Opt to use regular expressions for more complex search patterns. For instance, to find a **SSN** format, use `[0-9]{3}-[0-9]{2}-[0-9]{4}`. | No |
| **Case Sensitive?** | Determine if the search should be case sensitive. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave empty for all pages. The first page is numbered `0` (zero). Example: `0,1-2,5-`. | No |
| **Password** | If the **PDF** is password-protected, enter the password here. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `JSON Output`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AutoCropImages` | boolean | `false` | If you require to crop empty space around an inserted image use the following: `profiles": { 'AutoCropImages': true }` |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `YAdjustmentForReplacementText` | integer | - | Adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. See Adjust Text Alignment for more details. |
# Search and Replace With Image in PDF
Source: https://developer.pdf.co/integrations/make/search-and-replace-with-image
This unique feature enables you to search for text within a **PDF** document and replace it with an image. It’s particularly useful for enhancing documents with visual elements in place of specific text.
This unique feature enables you to search for text within a **PDF** document and replace it with an image. It's particularly useful for enhancing documents with visual elements in place of specific text.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Text to Search** | Specify the text that you wish to search for in the **PDF**. | No |
| **Replace Image Url** | Provide the URL of the image that will replace the found text. For example: `http://www.xyz.com/image.png` | No |
| **Use Regular Expressions** | Opt to use regular expressions for more complex search patterns. For instance, to find a **SSN** format, use `[0-9]{3}-[0-9]{2}-[0-9]{4}`. | No |
| **Case Sensitive?** | Determine if the search should be case sensitive. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave empty for all pages. The first page is numbered `0` (zero). Example: `0,1-2,5-`. | No |
| **Password** | If the **PDF** is password-protected, enter the password here. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `JSON Output`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `Page Count` | The total number of pages in the output **PDF**. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `AutoCropImages` | boolean | - | Controls whether to crop empty space around an inserted image. See Crop Empty Space Around Images for more information. |
# Search Text in PDF
Source: https://developer.pdf.co/integrations/make/search-text
This feature is designed to search for specific text within a **PDF** document. It is capable of searching text even in scanned **PDFs**, making it a versatile tool for document analysis and information retrieval.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File From URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File From URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Text to Search** | Specify the text that you wish to search for in the **PDF**. | Yes |
| **Use Regular Expressions** | Opt to use regular expressions for more complex search patterns. For instance, to find a **SSN** format, use `[0-9]{3}-[0-9]{2}-[0-9]{4}`. | No |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Pages** | Enter a comma-separated list of page indices (or ranges) for processing. Leave empty for all pages. The first page is numbered `0` (zero). Example: `0,1-2,5-`. | No |
| **Password** | If the **PDF** is password-protected, enter the password here. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `JSON Output`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `body` | An object array containing search results such as `text`, `left`, `top`, `width`, `height`, `pageIndex` and others. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| -------------------------------------------------- | ------- | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `ColumnDetectionMode` | string | `Content Groups And Borders` | Controls column detection/alignment in PDF table extraction. Modes: ContentGroupsAndBorders (default; text + lines), ContentGroups (text grouping only), Borders (lines only), BorderedTables (OCR-based for bordered tables), ContentGroupsAI (AI for dense/complex layouts). |
| `DetectionMinNumberOfRows` | integer | `1` | Minimum number of rows to detect in a table |
| `DetectionMinNumberOfColumns` | integer | `1` | Minimum number of columns to detect in a table |
| `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | `0` | Maximum number of invalid subsequent rows allowed in a table |
| `DetectionMinNumberOfLineBreaksBetweenTables` | integer | `0` | Minimum number of line breaks between tables |
| `EnhanceTableBorders` | boolean | `true` | Enhance table borders or not |
| `OCRDetectPageRotation` | boolean | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See Support page rotation for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `requestParametersDocument` | string | - | |
| `responseParameters` | object | | |
# Send Email With Attachments
Source: https://developer.pdf.co/integrations/make/send-email-with-attachments
This feature facilitates sending an email with attachments directly from your workflow. It’s ideal for automating email processes, including sending reports, invoices, or other documents as attachments.
## Input
| Name | Description | Required |
| ------------------ | --------------------------------------------------------------------------- | -------- |
| **Import Options** | Select the input source for attachments: `Upload Files` or `Input Link(s)`. | Yes |
***
**Upload Files**
| Name | Description | Required |
| -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
**Input Link(s)**
| Name | Description | Required |
| ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide URLs to source documents, separated by commas. Example: `https://example.com/sample1.pdf,https://example2.com/sample2.pdf`. Ensure links are publicly accessible if using services like **Google Drive** or **Dropbox**. | Yes |
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Email From** | Specify the sender's name and email. Example: `John Doe `. | Yes |
| **Email To** | Define the recipient's name and email. Example: `John Doe `. | Yes |
| **Subject** | The subject line of the email. | Yes |
| **Body Text** | The plain text version of the email message. | Yes |
| **Body HTML** | The HTML version of the email message. | Yes |
| **SMTP Server** | The SMTP server address. For setup details, refer to the [SMTP Configuration Guide](/knowledgebase/smtp-guide). | Yes |
| **SMTP Port** | The SMTP server port. | Yes |
| **SMTP User Name** | The SMTP server username. | Yes |
| **SMTP Password** | The SMTP server password. | Yes |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
## Output
| Name | Description |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{
'DataEncryptionAlgorithm': 'AES128',
'DataEncryptionKey': 'HelloThisKey1234',
'DataEncryptionIV': 'TreloThisKey1234'
}
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Split a PDF
Source: https://developer.pdf.co/integrations/make/split-pdf
This function allows you to split a single **PDF** file into multiple **PDF** files. It offers various methods for splitting, including by page index, page range, text search, or barcode search. This feature is particularly useful for segmenting large **PDF** documents or extracting specific sections.
## Input
| Name | Description | Required |
| ------------------ | ---------------------------------------------------------------------------- | -------- |
| **Import Options** | Choose the input source, either `Upload a File` or `Import a File from URL`. | Yes |
***
**Upload a File**
| Name | Description | Required |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload a file using raw binary data from another module. Note: This requires additional credits as it first uploads to [PDF.co Temporary Files Storage](/api/file-upload). | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Import a File from URL**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **URL** | Provide the URL to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | Specify a custom file name for the output file. | No |
**Split**
| Name | Description | Required |
| ------------ | -------------------------------------------------------------------------------- | -------- |
| **Split By** | Provide split by options. Valid values: `Page Numbers`, `Text Found or Barcode`. | Yes |
**Split by - Page Numbers**
| Name | Description | Required |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Pages** | Specify comma-separated page indices or ranges for splitting. The first page is numbered `1`. Example: `1,2,3-` or `1,2,3-7`. Use special notations like `!1` for the last page or `*` for splitting into separate pages. | Yes |
**Split by - Text Found or Barcode**
| Name | Description | Required |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Search String** | Specify the text or barcode to search for when splitting the **PDF**. Use macros for barcodes, such as `[[barcode:]]`. For example, `[[barcode:qrcode 12345]]` to search for a QR code with the value 12345. | No |
| **ExcludeKey Pages** | Choose to exclude pages where the search result is found. This option is useful if you want to omit certain pages based on the search criteria. | No |
| **Regex Search** | Enable this option if you are using regular expressions for your search. This allows for more complex search patterns. | No |
| **Case Sensitive** | Determine whether the search should be case sensitive. Enable this if the text case is important for your search criteria. | No |
| **Language** | Select the language used for **OCR** (Optical Character Recognition). This is relevant when searching text in scanned **PDFs** or images. | No |
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Execution Mode** | Select **Sync** for small tasks up to `10` seconds. Choose **Async** for standard jobs, or **Async For Large Docs** for tasks over `30` seconds. Use **Job Check** module for retrieving results in large tasks. | No |
| **Profiles** | Add custom options for the process in a `JSON` string format. See [API Profiles](#profiles) for more details. | No |
| **Output Links Expiration** | Set the expiration time in minutes for output links. Default is `60` minutes. Increase this limit with a `Business Plan` or higher, see [plans here](https://app.pdf.co/subscriptions) for details. | No |
| **Export Type** | Choose between `Download a File` or `JSON Output`. Default is `Download a File`. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `url` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `URLs` | Contains string array of split **PDF** URLs. This data is generated only when the `Export Type` option is set to `JSON Output`. |
| `Data` | Represents the output binary data. This data is generated only when the `Export Type` option is set to `Download a File`. |
| `Status` | Indicates the [response status](/api/introduction) code. A `success` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
### Profiles
To display the Profiles fields, you must **enable Advanced Settings** by clicking the toggle:
You can set additional options for the operation used in the [PDF.co](http://pdf.co/) module by using **Profiles**. A profile is a string in JSON-like format containing predefined parameters.
### Here’s an example of a Custom Profiles input:
```json theme={null}
{ "outputDataFormat": "base64" }
```
With this input, the [PDF.co](http://pdf.co/) module will return the output in base64 format. You can find the list of available parameters for customizing profiles in the [PDF.co](http://pdf.co/) operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within Make using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Upload a File
Source: https://developer.pdf.co/integrations/make/upload-file
This function allows for uploading a file to the secure [PDF.co Temporary Files Storage](/api/file-upload). Uploaded files are stored for one hour and can be used in conjunction with other modules and applications during this period.
## Input
| Name | Description | Required |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Data** | Upload the source file to [PDF.co Temporary Files Storage](/api/file-upload). The uploaded file will be available for use within 1 hour. The generated URL can be used with other **PDF.co** modules and apps. | Yes |
| **File Name** | Optionally specify a custom file name for the uploaded file. This should be a string. | No |
### Integrating External File Sources
Streamline your **Make** workflows with external file sources like **Google Drive** and **Dropbox** using their unique actions. Discover efficient integration strategies in our guide: [File Source Integrations in Make](/integrations/make/input-file-sources).
***
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------------------------------------------- |
| `URL` | This is the temporary **URL** provided by the **PDF.co** file server. |
| `Status` | Indicates the [response status](/api/introduction) code. A `200` status is returned if the operation is successful. |
| `outputLinkValidTill` | Specifies the timestamp until which the `url` remains accessible. |
| `error` | Provides details about any errors encountered during the process, if applicable. |
| `File Name` | The designated name of the output file. |
| `Job Id` | A unique identifier assigned to the job. |
| `credits` | The amount of credits utilized for the process. |
| `Remaining Credits` | Displays the balance of credits available in your account. |
| `duration` | The duration of time the process took to complete. |
## Bypassing Google Drive’s Temporary Blocks
Google Drive may temporarily block PDF.co’s access to publicly shared URLs, identifying it as bot activity. This blockage prevents PDF.co from accessing the input file, causing errors in Make scenarios.
**Solution**: Use the “Upload a File” module to upload your file directly to PDF.co. This will provide a PDF.co URL for your file, which you can use in subsequent steps without encountering access issues.
## Steps:
* **Download the File**: First, download your file from Google Drive.
* **Upload to PDF.co**: Use PDF.co’s “Upload a File” module to upload the downloaded file.
* **Use PDF.co URL**: You will receive a PDF.co-hosted URL for your file. Use this URL in your Make scenario steps to avoid Google Drive’s temporary blocks.
This method ensures uninterrupted access to your files for processing with PDF.co in your Make scenarios.
# Getting Started with Microsoft Power Automate
Source: https://developer.pdf.co/integrations/microsoft-power-automate/getting-started
[Microsoft Power Automate](https://make.powerautomate.com) formerly known as **Microsoft Flow** is one of the main features of **Microsoft’s Power Platform**. It is a service that can be used to automate business workflow. It has a guided and no\-code functionality that makes it possible for everyone to easily automate time\-consuming tasks and processes. It offers seamless integration to other apps using a pre\-built **Connector Library** to create small and large\-scale systems.
## Microsoft Power Automate & PDF.co
Need help and support? Contact us using the button below:
## How to Convert HTML to PDF
# Integrating File Sources with PDF.co
Source: https://developer.pdf.co/integrations/microsoft-power-automate/input-file-sources
Integrating SharePoint, OneDrive, and other storage platforms into your Power Automate workflows with PDF.co can streamline document processing. This guide provides detailed steps for linking these sources and efficiently managing documents with PDF.co in Power Automate.
## URL Accessibility for PDF.co
The API supports TLS 1.2 and 1.3 for secure connections. Earlier versions such as TLS 1.0 and 1.1 are deprecated and should be avoided.
To integrate files from SharePoint, OneDrive, and similar services with PDF.co, URLs must be accessible. Power Automate’s file-sharing links are recommended, as they are PDF.co-compatible. For direct URLs outside Power Automate, ensure public accessibility so PDF.co can access files without restrictions.
## Integrating SharePoint Files with PDF.co in Power Automate
You can automate document processing by generating a direct link for SharePoint files, enabling PDF.co to process them and save outputs back to SharePoint.
**Steps to Generate the Direct Download Link:**
1. **Add the “Create sharing link for a file or folder” Action**
* In Power Automate, add SharePoint’s **“Create sharing link for a file or folder”** action to create a shareable link that can be adjusted for direct access.
* **Site Address**: Select your SharePoint site.
* **Library Name**: Choose the library containing your file.
* **Item Id**: Specify the file you want to share.
For more information, see [Microsoft Docs on creating sharing links](https://docs.microsoft.com/en-us/connectors/sharepointonline/#create-sharing-link-for-a-file-or-folder).
2. **Modify the Link for PDF.co Compatibility**
* Append `?download=1` to convert it into a direct download link compatible with PDF.co.
**Example:**
Original Link:
```javascript theme={null}
https://yourcompany.sharepoint.com/:b:/s/SharedDocuments/Ed1...4ig
```
Modified Link:
```javascript theme={null}
https://yourcompany.sharepoint.com/:b:/s/SharedDocuments/Ed1...4ig?download=1
```
3. **Troubleshooting**:
* Ensure the sharing settings allow public access.
* Test the link in a browser to confirm it downloads the file directly.
* Double-check permissions if PDF.co cannot access the file.
## Storing PDF.co Output to SharePoint
After processing files with PDF.co, the output can be stored in SharePoint for easy access and further use.
**Steps to Save PDF.co Output:**
1. **Retrieve PDF.co Output Using the HTTP Action**
* Add an **HTTP** action to download the processed file from PDF.co.
* **Method**: Set to `GET`.
* **URI**: Enter the **Output URL** provided by PDF.co after processing.
* This action will fetch the processed file content for saving in SharePoint.
2. **Create a File in SharePoint**
* Add the **“Create file”** action in SharePoint to store the retrieved file.
* **Site Address**: Choose your SharePoint site.
* **Folder Path**: Choose the destination folder within your document library.
* **File Name**: Assign a name to the file, including the extension (e.g., `ProcessedFile.pdf`).
* **File Content**: Use the **Body** output from the HTTP action.
Additional guidance is available in [Microsoft Docs on creating files](https://docs.microsoft.com/en-us/connectors/sharepointonline/#create-file).
By following these steps, you can fully integrate SharePoint files with PDF.co in Power Automate, streamlining your document processes for improved productivity.
Additional Resources:
* [Download image from url and save to sharepoint as image](https://community.powerplatform.com/forums/thread/details/?threadid=cd5daea6-cfcf-4395-88f3-191928a6c247)
* [Image attachments not displaying when uploaded to SharePoint](https://community.powerplatform.com/forums/thread/details/?threadid=7f088f1e-cf62-429f-9f6c-ab628bd682dd)
* [How to upload files from Public URLs to SharePoint using PowerAutomate](https://sharepoint.handsontek.net/2024/02/20/upload-files-public-urls-sharepoint-using-powerautomate/).
# Add Text or Images to PDF
Source: https://developer.pdf.co/integrations/n8n/add-text-image-to-pdf
Add text, images, and annotations to your PDF documents with precise positioning and styling options.
## Input
| Name | Description | Required |
| :-------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text** | Enter the text for the text object. Macros like line breaks (`\\n` or `{{$$newLine}}`) or page numbers (`{{$$PageNumber}}`) can be inserted. For using special macros on **n8n**, see [Macros for Text](/knowledgebase/macros-for-text#special-macro-style%3A-square-brackets). | Yes |
| **Image URL** | URL to image, base64 encoded image, or `filetoken://` link to image stored in [PDF.co Files storage](https://app.pdf.co/tools/files). | Yes |
| **X Coordinate** | Determine the X coordinate for text placement. Use **PDF.co** [**PDF Inspector**](https://app.pdf.co/pdf-inspector) to find or measure **PDF** **coordinates**. | Yes |
| **Y Coordinate** | Specify the Y coordinate. Use [****PDF.co PDF Inspector****](https://app.pdf.co/pdf-inspector) to find or measure **PDF coordinates**. | Yes |
| **Font Size** | Set the font size, with the default being `12`. | No |
| **Font Color** | Choose the text color in hex format (`#RRGGBB` or `#AARRGGBB`, with `AA` as transparency). The default is `#000000`. | No |
| **Font Bold** | Enable this to apply bold styling to the font. | No |
| **Font Italic** | Enable this to apply italic styling to the font. | No |
| **Font Strikeout** | Enable this to apply strikeout styling to the font. | No |
| **Font Underline** | Enable this to apply underline styling to the font. | No |
| **Font Name** | Select the font name from the [**available font list**](https://developer.pdf.co/knowledgebase/general). The default font is **Arial**. | No |
| **Pages** | Indicate specific page numbers or ranges where the text should be added. Leave blank to include all pages. The first page is numbered `0`. Example: `0,2-5,7-`. | No |
| **Link** | Add an optional clickable link (starting with `http://`, `https://`, `mailto:name@example.com`, etc.). | No |
| **Width** | Define the width of the text box. Coordinates start at the top left (use the [**PDF.co PDF Inspector**](https://app.pdf.co/pdf-inspector) to measure coordinates). | No |
| **Height** | Define the height of the text box. Coordinates start at the top left (use the [**PDF.co PDF Inspector**](https://app.pdf.co/pdf-inspector) to measure coordinates). | No |
| **Alignment** | Set text alignment as `Center`, `Right`, or `Left`. Default is `Center`. | No |
| **Transparent** | Set the text background as transparent. | No |
| **Keep Aspect Ratio** | Keep the aspect ratio of the image. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api)within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :------------- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew. |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
You can also use `Custom Profiles` to:
* [Crop a PDF File](https://developer.pdf.co/api/pdf-add#crop-a-pdf-file)
* [Disable Ligaturization](https://developer.pdf.co/api/pdf-add#disable-ligaturization)
* [Flatten Document](https://developer.pdf.co/api/pdf-add#flatten-document)
* Visit this page for [general information on Profiles usage.](https://developer.pdf.co/api/profiles)
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# AI Invoice Parser
Source: https://developer.pdf.co/integrations/n8n/ai-invoice-parser
Automate invoice data extraction with PDF.co’s AI-powered invoice parser. Instantly convert PDFs into structured JSON by detecting key fields like invoice number, vendor name, amounts, dates, and itemized lines — no manual setup or templates required.
**Important**
* **Only invoice documents** are supported for parsing.
* To ensure accurate processing, each invoice must be clearly separated. **If an invoice contains multiple pages, we recommend splitting it** into individual PDFs using the [PDF Split API](/integrations/n8n/split-pdf).
* While AI Invoice Parser supports multi-page invoices, **the total page count for a single PDF must not exceed 100 pages**. Submitting large PDFs containing multiple invoices is not recommended.
## Input
| Name | Description | Required |
| :-------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **Invoice File URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Custom Field** | Specify custom fields beyond the default list. Use comma-separated format for multiple custom fields extraction (e.g., `storeNumber`, `lineTotal`, `financialCharges`). | No |
| **lineItemStructure** | A JSON object that defines a [custom structure](/integrations/n8n/ai-invoice-parser#line-item-structure) for line items. Each key is a field name (in `camelCase`) and each value is `"string"` or `"number"`. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
## File Requirements and Limitations
The PDF.co AI Invoice Parser only supports files with **up to 100 pages**. Files exceeding this limit will **not be processed**, so please avoid uploading documents larger than 100 pages.
If your file contains **multiple invoices in a single PDF**, use the [Split PDF](/integrations/n8n/split-pdf) operation in n8n to separate each invoice before passing them to the AI Invoice Parser node.
## Source PDF URLs
If your previous node or trigger outputs a binary file (e.g., from Gmail or other apps), you’ll need to upload the file first using the PDF.co [Upload File operation](/integrations/n8n/upload-file) to generate a **valid URL**. This URL is required as the input for the Source File URL field in all subsequent PDF.co actions. You can integrate with **Google Drive**, **Dropbox**, or other apps to trigger your automation and pass files into PDF.co.
## Custom Field Extraction
AI Invoice Parser with custom fields support automatically detects invoice layouts and extracts both standard schema data and user-specified custom fields without requiring manual templates.
The customField parameter allows you to specify additional fields to extract beyond the standard schema. Some examples include:
* `storeNumber` – Store or branch identifier
* `deliveryDate` – Expected delivery date
* `financialCharges` – Additional financial charges
* `lineTotal` – Total amount for line items
* `purchaseOrderRef` – Purchase order reference number
* `customerReference` – Customer reference number
* `departmentCode` – Department or cost center code
*If a custom field returns an empty value, please* [contact our support team](https://pdf.co/support/request?subject=ai-invoice-parser%20-%20custom%20fields) *to help improve the extraction accuracy.*
## Line Item Structure
The `lineItemStructure` parameter lets you define a custom schema for line items. Each key is a field name you choose (in `camelCase`) and each value is the expected data type — either `"string"` or `"number"`.
When provided, every object inside the `lineItems` array will contain exactly the fields you specified. If a value cannot be extracted from the invoice, the field will still be present with an empty or default value instead of being omitted.
Use `camelCase` for field names (e.g., `unitPrice`, `totalPrice`). The field names you define will be used as-is in the response, giving you full control over the output keys.
## Output
| Name | Description |
| :----------------- | :----------------------------------------------------------------------------------- |
| `PageCount` | Total Page Count. |
| `url` | Direct URL to the final PDF file stored in S3. |
| `body` | An object array containing the all invoice parsing result. |
| `duration` | The time it took for the process. |
| `error` | Details of any errors (if any). |
| `status` | The [**response status**](/api/response-codes) code. If all good this will be `200`. |
| `jobId` | Unique identifier for the background job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
## Supported Languages
* **Albanian (Shqip)**
* **Bosnian (Bosanski)**
* **Bulgarian (Български)**
* **Croatian (Hrvatski)**
* **Czech (Čeština)**
* **Danish (Dansk)**
* **Dutch (Nederlands)**
* **English**
* **Estonian (Eesti)**
* **Finnish (Suomi)**
* **French (Français)**
* **German (Deutsch)**
* **Greek (Ελληνικά)**
* **Hungarian (Magyar)**
* **Icelandic (Íslenska)**
* **Italian (Italiano)**
* **Latvian (Latviešu)**
* **Lithuanian (Lietuvių)**
* **Norwegian (Norsk)**
* **Polish (Polski)**
* **Portuguese (Português)**
* **Romanian (Română)**
* **Russian (Русский)**
* **Serbian (Српски)**
* **Slovak (Slovenčina)**
* **Slovenian (Slovenščina)**
* **Spanish (Español)**
* **Swedish (Svenska)**
* **Turkish (Türkçe)**
* **Ukrainian (Українська)**
# Barcode Generator
Source: https://developer.pdf.co/integrations/n8n/barcode-generator
Generate various types of barcodes including QR Code, Code 128, Code 39, and more.
## Input
| Name | Description | Required |
| :---------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **Barcode Value** | Set the string value to encode inside the barcode, must be in a string format. | Yes |
| **Barcode Type** | Set the barcode type to be used. | Yes |
| **Decoration Image** | Enter the image URL to insert as a logo inside the QR Code. Only **PNG**, **JPG**, or **JPEG** formats are supported. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Generate Inline URL** | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api)within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :------ | :-------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `Angle` | integer | `0` | See [**profiles.Angle**](https://developer.pdf.co/api/barcode/generate#profiles-angle) |
| `NarrowBarWidth` | integer | 3 | See [**profiles.NarrowBarWidth**](https://developer.pdf.co/api/barcode/generate#profiles-narrowbarwidth) |
| `CaptionFont` | string | Arial, 12 | See [**profiles.CaptionFont**](https://developer.pdf.co/api/barcode/generate#profiles-captionfont) |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](https://developer.pdf.co/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Barcode Reader
Source: https://developer.pdf.co/integrations/n8n/barcode-reader
Decode barcodes from images or PDF documents quickly and accurately.
## Input
| Name | Description | Required |
| :--------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Barcode Type** | Choose the type of barcode to read. By default, the system will look for QR Codes. | Yes |
| **Pages** | Specify page indices as comma-separated values or ranges to process (e.g. “`0, 1, 2-`” or “`1, 2, 3-7`”). The first-page index is `0`. Use ”`!`” before a number for inverted page numbers (e.g. “`!0`” for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Optical Marks Reader** | Comma-separated list of additional marks to detect. The barcode reader engine can also find marks like **Checkbox**, **UnderlinedField**, etc. on scanned documents. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{
'DataEncryptionAlgorithm': 'AES128',
'DataEncryptionKey': 'HelloThisKey1234',
'DataEncryptionIV': 'TreloThisKey1234'
}
```
With this input, the PDF.co operation will encrypt the output with strong `AES128` encryption. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](https://developer.pdf.co/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Compress PDF
Source: https://developer.pdf.co/integrations/n8n/compress-pdf
This operation allows you to compress/reduce the size of your PDF files while maintaining the quality.
## Input
| Name | Description | Required |
| :------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| PDF URL | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| File Name | File name for the generated output, the input must be in string format. | No |
| Webhook URL | The callback URL or Webhook used to receive the output data. | No |
| Password | The password of the password-protected PDF file | No |
| HTTP Username | HTTP auth user name if required to access source URL. | No |
| HTTP Password | HTTP auth password if required to access source URL. | No |
| Custom Compression Configuration | [**Configuration options**](https://developer.pdf.co/api/pdf-compress#config). By default this object is pre-defined to give typical options for optimization. | No |
## Custom Compression Configuration
The `config` object defines granular settings for the compression process. By default, this `config` object is set to match the standard configuration for image optimization by Adobe Acrobat Pro and is defined as follows:
```json theme={null}
{
"images": {
"color": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 150,
"threshold_ppi": 225
},
"compression": {
"skip": false,
"compression_format": "jpeg",
"compression_params": {
"quality": 60
}
}
},
"grayscale": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 150,
"threshold_ppi": 225
},
"compression": {
"skip": false,
"compression_format": "jpeg",
"compression_params": {
"quality": 60
}
}
},
"monochrome": {
"skip": false,
"downsample": {
"skip": false,
"downsample_ppi": 300,
"threshold_ppi": 450
},
"compression": {
"skip": false,
"compression_format": "ccitt_g4",
"compression_params": {}
}
}
},
"save": {
"garbage": 4
}
}
```
Refer to this [guide for more details](https://developer.pdf.co/api/pdf-compress#config) on the config object.
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Convert PDF to Anything
Source: https://developer.pdf.co/integrations/n8n/convert-from-pdf
Convert your PDF files into various document formats such as CSV, HTML, XML, JSON, TXT, images (JPG, PNG, TIFF, and WEBP), and spreadsheets (XLS or XLSX).
## Input
| Name | Description | Required |
| :-------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Conversion Type** | Choose which document format you want to convert your PDF to, such as CSV, HTML, TXT, JSON, XML, JPG, PNG, TIFF, WEBP, XLS or XLSX. | Yes |
| **Pages** | Specify page indices as comma-separated values or ranges to process (e.g. “0, 1, 2-” or “1, 2, 3-7”). The first-page index is 0. Use ”!” before a number for inverted page numbers (e.g. “!0” for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | No |
| **Line Grouping** | Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: `1`, `2`, `3`. For more information, see [Line Grouping](https://developer.pdf.co/api/pdf-to-csv#line-grouping-options). | No |
| **Unwrap** | Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when `lineGrouping` is set to `1`. | No |
| **OCR Language Name or ID** | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | No |
| **Extraction Region** | Defines coordinates for extraction. Use [PDF Edit Add Helper](https://app.pdf.co/pdf-edit-add-helper) to get or measure PDF coordinates. The format is `{x} {y} {width} {height}`. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :----------------------------- | :---------------------------- | :------------------------ | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64`.
This profile parameter is only available for PDF to **JPG**, **PNG**, **WEBP**, and **TIFF** operations. |
| `ColumnDetectionMode` | string | `ContentGroupsAndBorders` | Controls column detection/alignment in PDF table extraction. See [**Column Detection Mode**](https://developer.pdf.co/api/pdf-to-csv#column-detection-mode) for more information.
This profile parameter is only available for PDF to **CSV** and **XLS** operations. |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [**OCR Extraction Modes**](https://developer.pdf.co/api/profiles#ocr-extraction-modes).
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator).
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to `true` to consider font colors.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to `true` to automatically align columns to the header row. When set to `true` (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to `false`, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Use methods like `AddGammaCorrection()` and `AddGrayscale()` by including them as JSON keys. [See sample usage](https://developer.pdf.co/api/pdf-to-image/png#ocrimagepreprocessingfilters). |
| `.AddGammaCorrection()` | array\[string (float format)] | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `.AddGrayscale()` | boolean | `false` | Set to `true` to preprocessing filter that converts a colored document/image to grayscale before performing OCR.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `RenderTextObjects` | boolean | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See [Disable Text Layer](https://developer.pdf.co/api/pdf-to-image/jpg#disable-text-layer) for more information.
This profile parameter is only available for PDF to **JPG, PNG, WEBP** and **TIFF** operations. |
| `RenderImageObjects` | boolean | `true` | Render image objects or not.
This profile parameter is only available for PDF to **JPG, PNG** and **WEBP** operations. |
| `RenderVectorObjects` | boolean | `true` | Render vector objects or not.
This profile parameter is only available for PDF to **JPG, PNG** and **WEBP** operations. |
| `JPEGQuality` | integer | `85` | See [profiles.JPEGQuality](https://developer.pdf.co/api/pdf-to-image/jpg#jpegquality)
This profile parameter is only available for PDF to **JPG** operation. |
| `WEBPQuality` | integer | `75` | See [profiles.WEBPQuality](https://developer.pdf.co/api/pdf-to-image/webp#profiles-webpquality)
This profile parameter is only available for PDF to **WEBP** operation. |
| `TIFFCompression` | string | `LZW` | See [profiles.TIFFCompression](https://developer.pdf.co/api/pdf-to-image/tiff#profiles-tiffcompression)
This profile parameter is only available for PDF to **TIFF** operation. |
| `RenderingResolution` | integer | `120` | See [Set Image Resolution](https://developer.pdf.co/api/pdf-to-image/jpg#set-image-resolution) for more information.
This profile parameter is only available for PDF to **JPG, PNG, WEBP** and **TIFF** operations. |
| `OptimizeImages` | boolean | `true` | Some PDF may have high quality images used in the document and you may need to keep the quality of these images in the output HTML. By default PDF to HTML is optimizing images and you can easily turn it off. See [Control Image Quality](https://developer.pdf.co/api/pdf-to-html#control-image-quality) for more information.
This profile parameter is only available for PDF to **HTML** operation. |
| `OutputPageWidth` | integer | `1024` | Control page width (in pixels) for output HTML. Height is calculated and used according to the original pdf pages ratio. See [Control Output Page Width](https://developer.pdf.co/api/pdf-to-html#control-output-page-width) for more information.
This profile parameter is only available for PDF to **HTML** operation. |
| `AdditionalCssStyles` | string | `“` | To inject CSS for layout options in your HTML. Example: `#canvas { zoom: 50%; }`. Scale the div that contains all generated HTML pages by 50%. See [Inject CSS](https://developer.pdf.co/api/pdf-to-html#inject-css) for more information.
This profile parameter is only available for PDF to **HTML** operation. |
| `SaveVectors` | boolean | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to `true` to save vector graphics.
This profile parameter is only available for PDF to **CSV, JSON, XLS** and **XML** operations. |
| `SaveImages` | string | `None` | Controls how images are saved during PDF to HTML conversion. Modes: `None` (no images), `OuterFile` (save to sub-folder), `Embed` (embed as Base64 data:URI).
This profile parameter is only available for PDF to **CSV, JSON, XLS, XML** and **HTML** operations. |
| `ConsiderFontSizes` | boolean | `false` | Set to `true` to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements.
This profile parameter is only available for PDF to **CSV, JSON, XLS** and **XML** operations. |
| `ExtractionArea` | array\[number] | - | Extract text in a specific area by defining the extraction area - set with points in the format `[x, y, width, height]`.
This profile parameter is only available for PDF to **CSV, JSON, XLS** and **XML** operations. |
| `ExtractShadowLikeText` | boolean | `true` | Controls whether to extract invisible text from a PDF document. Set to `false` to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to `Auto` to properly apply the shadow text filtering effect.
This profile parameter is only available for PDF to **CSV, JSON, XLS** and **XML** operations. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information.
This profile parameter is available for PDF to **CSV, JSON, Text, XLS, XML, HTML, JPG, PNG, WEBP** and **TIFF** operations. |
You can also use `Custom Profiles` to:
* Fix incorrect data positions caused by overlapping invisible text or objects. For more details, please refer to this [guideline](https://developer.pdf.co/api/pdf-to-csv#column-detection-mode).
* [Disable images.](https://developer.pdf.co/api/pdf-to-html#disable-images)
* [Control image quality.](https://developer.pdf.co/api/pdf-to-html#control-image-quality)
* [Control output page width.](https://developer.pdf.co/api/pdf-to-html#control-output-page-width)
* [Inject CSS](https://developer.pdf.co/api/pdf-to-html#inject-css).
* [Disable text layer](https://developer.pdf.co/api/pdf-to-image/png#disable-text-layer).
* [Set image resolution](https://developer.pdf.co/api/pdf-to-image/png#set-image-resolution).
* [Set image preprocessing filters.](https://developer.pdf.co/api/pdf-to-image/png#ocrimagepreprocessingfilters)
* Visit this page for [general information on Profiles usage.](https://developer.pdf.co/api/profiles)
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Convert Anything to PDF
Source: https://developer.pdf.co/integrations/n8n/convert-to-pdf
This operation allows you to convert a variety of file types into PDF format. It supports transforming CSV, document (RTF, DOC, DOCX, TXT), email (MSG or EML), image (JPG, PNG, and TIFF), and spreadsheet (XLS or XLSX) into high-quality PDF files easily.
## Input
| Name | Description | Required |
| :---------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Conversion Type** | Choose from which document you want to convert to PDF, such as **CSV**, **Document** (RTF, DOC, DOCX, TXT), **Email** (MSG or EML), **Image** (JPG, PNG, TIFF), or **Spreadsheet** (XLS and XLSX). | Yes |
| **Auto Size** | Set to true to adjust page dimensions to content with automatic page sizing. If `false`, uses worksheet’s page setup. | No |
| **Embed Attachment** | Set to `false` if you don’t want to convert attachments from the original email and want to embed them as original files (as embedded PDF attachments). Converts attachments that are supported by the PDF.co API (DOC, DOCx, HTML, PNG, JPG etc.) into PDF format and then merges into output final PDF. Non-supported file types are added as PDF attachments (Adobe Reader or another viewer may be required to view PDF attachments). | No |
| **Convert Attachment to PDF** | Set to `true` to automatically embeds all attachments from original input email `MSG` or `EML` files into the final output PDF. Set it to `false` if you don’t want to embed attachments so it will convert only the body of the input email. True by default. | No |
| **Margins** | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | No |
| **Orientation** | Set the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | No |
| **Paper Size** | Specifies the paper size. Accepts standard sizes like `Letter`, `Legal`,`Tabloid` , `Ledger`, `A0` - `A6`. | No |
| **Custom Paper Size** | You can custom paper size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: ‘`200 300`’, ‘`200px 300px`’, ‘`200mm 300mm`’, ‘`20cm 30cm`’, ‘`6in 8in`’. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# PDF.co API and n8n Integration Guide
Source: https://developer.pdf.co/integrations/n8n/custom-api-call
PDF.co provides various document files processing like extraction, editing, conversion and more. You can integrate them with your n8n workflow by using the n8n HTTP Request node to leverage the full power of document automation.
## Comprehensive API Capabilities
### **1. Document Intelligence & Parsing**
* **AI-Powered Document Parser**: Extract structured data from invoices, receipts, forms, and custom documents
* **Document Classifier**: Automatically categorize and route documents
* **AI Invoice Parser**: Specialized invoice data extraction with high accuracy
* **Optical Mark Recognition**: Read checkboxes, radio buttons, and fillable fields
### **2. PDF Manipulation & Processing**
* **PDF Split & Merge**: Combine multiple PDFs or split by pages, text search, or barcode detection
* **PDF Form Operations**: Fill forms, create fillable forms, and extract form data
* **Text Operations**: Search, replace, add, or delete text with precision
* **Image Integration**: Add signatures, images, and replace text with images
* **Page Management**: Rotate, delete, and reorganize PDF pages
* **Security Features**: Password protection and user-controlled encryption
### **3. Format Conversion Excellence**
* **From PDF**: Convert to CSV, JSON, TEXT, XLS, XLSX, XML, HTML, JPG, PNG, WEBP, TIFF
* **To PDF**: Generate from HTML, URL, images, CSV, XLS, Word documents, and HTML templates
* **Spreadsheet Processing**: Convert XLS/XLSX to CSV, JSON, HTML, TXT, XML
### **4. Barcode & QR Code Support**
* **Barcode Generation**: Create various barcode formats programmatically
* **Barcode Reading**: Extract data from existing barcodes in documents
### 5. **Email & Communication**
* **Email Integration**: Send and decode emails, convert emails to PDF
* **PDF from Email**: Generate PDFs from email content
## How to Call PDF.co API on n8n
1. **Add an **[**HTTP Request Node**](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.httprequest/)** to your n8n workflow**
2. **Input the API method and URL**
You can find all of our API endpoint URLs on our [doc page](https://developer.pdf.co/api). Make sure you add `https://api.pdf.co` before inputting the URL shown on our doc page. Here's the sample URL input:
```
```
3. Use `Predefined Credential Type` for Authentication and choose PDF.co API from the selection menu.
4. Add your PDF.co credentials by inputting your PDF.co API Key, which you can get from the [PDF.co dashboard](https://app.pdf.co).
5. Toggle the Send Body switch to `ON` and select **JSON** from the dropdown.
6. You can input the body parameters by using the parameter field or input directly as **JSON** instead (see the sample below).
```
{
"url": "",
"callback": ""
}
```
# Delete PDF Pages
Source: https://developer.pdf.co/integrations/n8n/delete-page-in-pdf
This operation allows you to delete one or more pages from your PDF document.
## Input
| Name | Description | Required |
| :--------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Pages** | Specify the pages you want to delete using comma-separated values or page ranges (e.g., “`1,2,3-`” or “`1,2,3-7`”). The first page is `1`. Inverted (“`!`”) page numbers are not supported by this operation. | Yes |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like format** containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Fill a PDF Form
Source: https://developer.pdf.co/integrations/n8n/fill-a-pdf-form
Fill interactive PDF form with text, checkboxes, radio buttons and other form elements.
## Input
| Name | Description | Required |
| :--------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Form Field Name** | Name of the form field. To find form fields please use the PDF Information operation or our [PDF Edit Add Helper tool](https://app.pdf.co/pdf-edit-add-helper). | Yes |
| **Text** | Enter the text you want to insert into the form field. If you need to check a checkbox field then set to `true`. For radio box, set index like `1`. | No |
| **Pages** | Indicate specific page numbers or ranges where the text should be added. Leave blank to include all pages. The first page is numbered `0`. Example: `0,2-5,7-`. | No |
| **Font Size** | Set the font size, with the default being `12`. | No |
| **Font Color** | Choose the text color in hex format (`#RRGGBB` or `#AARRGGBB`, with `AA` as transparency). The default is `#000000`. | No |
| **Font Bold** | Enable this to apply bold styling to the font. | No |
| **Font Italic** | Enable this to apply italic styling to the font. | No |
| **Font Strikeout** | Enable this to apply strikeout styling to the font. | No |
| **Font Underline** | Enable this to apply underline styling to the font. | No |
| **Font Name** | Select the font name from the [**available font list**](/knowledgebase/general). The default font is **Arial**. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :------------- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle `[x, y, width, height]` in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew. |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Get PDF Information & Form Fields
Source: https://developer.pdf.co/integrations/n8n/get-pdf-information
Extract information from a PDF document, including form fields, page count, size, author, description, keywords, and more.
## Input
| Name | Description | Required |
| :--------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Extract Fillable Fields** | Enable this option to extract information from fillable form fields in the PDF. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{'OCRMode': 'TextFromImagesAndVectorsAndRepairedFonts'}
```
With this input, the PDF.co operation will extract text from images and repaired fonts from documents. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :----------------------------- | :---------------------------- | :-------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [**OCR Extraction Modes**](https://developer.pdf.co/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to `true` to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to `true` to automatically align columns to the header row. When set to `true` (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to `false`, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Use methods like `AddGammaCorrection()` and `AddGrayscale()` by including them as JSON keys. [See sample usage](https://developer.pdf.co/api/pdf-to-image/png#ocrimagepreprocessingfilters). |
| `.AddGammaCorrection()` | array\[string (float format)] | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. |
| `.AddGrayscale()` | boolean | `false` | Set to `true` to preprocessing filter that converts a colored document/image to grayscale before performing OCR. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Getting Started with n8n
Source: https://developer.pdf.co/integrations/n8n/getting-started
The PDF.co n8n community node enables fast and easy automation of PDF-related tasks within n8n workflows. Built on the PDF.co API, this node allows you to programmatically convert, merge, split, and extract data from PDF documents. As a community-maintained integration, it helps streamline document processing without complex coding, making it ideal for automating business operations involving PDFs.
## Installation
### Installing PDF.co Node on N8N (Version 1.94.0 and Later)
1. **Open the Nodes Panel**\
In your N8N instance, navigate to the nodes panel.
2. **Search for the “PDF.co” Node**\
In the search bar, type "pdf.co" to find the PDF.co node.
3. **Install the Node**\
Click on the **PDF.co Api** node, and in the node details window, click the **Install Node** button.
## Operations
This node provides comprehensive PDF processing capabilities through PDF.co's API. Here are the available features:
### **1. AI-Powered Document Processing**
* AI Invoice Parser: Extract data from invoices using AI-powered parsing
### **2. URL/HTML to PDF Conversion**
* URL to PDF
* HTML to PDF
* HTML Template to PDF
### **3. PDF Merging**
* Merge multiple PDFs into one
* Support for merging from different file formats (merge2)
### **4. PDF Splitting**
* Split by page number
* Split by search text
* Split by barcode
### **5. Document to PDF Conversion**
* Document to PDF (RTF, DOC, DOCX, TXT)
* Spreadsheet to PDF (CSV, XLS, XLSX, TXT Spreadsheet)
* Image to PDF (JPG, PNG, TIFF)
* Email to PDF (MSG or EML)
### **6. PDF to Other Formats**
* PDF to CSV
* PDF to HTML
* PDF to Images (JPG, PNG, TIFF, WEBP)
* PDF to JSON
* PDF to Text
* PDF to Excel (XLS/XLSX)
* PDF to XML
### **7. PDF Modification**
* Add text annotations
* Add images to PDF documents
* Fill interactive PDF forms with data
* Extract PDF metadata
* Get form field information
### **8. PDF Optimization**
* Compress PDF files
* Optimize PDF for web or storage
* Remove password protection
* Add password protection
* Rotate pages in PDF documents
* Remove specific pages from PDF
### **9. PDF Search and Text Operations**
* Search for text within PDF documents
* Search and delete text
* Search and replace text
* Search and replace with image
### **10. Barcode Operations**
* Extract barcode data from PDFs
* Generate barcodes in PDFs
### **11. PDF OCR and Searchability**
* Make scanned PDF searchable (OCR)
* Make searchable PDF unsearchable
### **12. File Management**
* Standard file upload
* Create file URL from input text/content
* Upload from Base64 encoded data
## Credentials
To use this node, you need a PDF.co API key. Here's how to get started:
1. Sign up for a PDF.co account at [**https://pdf.co**](https://pdf.co)
2. Navigate to your dashboard and obtain your API key
3. In n8n, add your PDF.co credentials by providing your API key
## Usage
This node allows you to automate PDF processing tasks in your n8n workflows. Here are some common use cases:
* Process invoices using AI-powered parsing
* Automatically convert HTML reports to PDF
* Merge multiple PDF documents into a single file
* Extract text from PDF documents for data processing
* Add watermarks to PDF documents
* Compress PDF files to reduce size
* Secure PDFs with password protection
* Convert various document formats to PDF
* Extract data from PDFs to other formats
* Manage and modify PDF documents programmatically
# Make PDF Searchable/Unsearchable
Source: https://developer.pdf.co/integrations/n8n/make-pdf-searchable-or-unsearchable
Transform scanned PDFs or images into fully searchable documents with OCR technology. This operation also allows you to make your PDF files unsearchable to protect sensitive information.
## Input
| Name | Description | Required |
| :--------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Make PDF Searchable/Unsearchable** | Choose whether you want to make the PDF searchable or non-searchable. | Yes |
| **OCR Language Name or ID** | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see [Language Support](https://developer.pdf.co/api/language-support). You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | Yes |
| **Pages** | Specify page indices as comma-separated values or ranges to process (e.g. “0, 1, 2-” or “1, 2, 3-7”). The first-page index is 0. Use ”!” before a number for inverted page numbers (e.g. “!0” for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **Password** | The password of the password-protected PDF file | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profile](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :----------------------------- | :---------------------------- | :-------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `OCRMode` | string | Auto | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see [**OCR Extraction Modes**](https://developer.pdf.co/api/profiles#ocr-extraction-modes). |
| `OCRResolution` | integer | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from `72` to `1200` dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: `0`, `1`, `2`, `3`. |
| `LineGroupingMode` | string | None | Controls line grouping in PDF text extraction. Modes: `None` (no grouping), `GroupByRows` (merge rows if all cells align), `GroupByColumns` (merge cells by column), `JoinOrphanedRows` (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to `true` to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | 1.2 | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to `true` to automatically align columns to the header row. When set to `true` (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to `false`, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Use methods like `AddGammaCorrection()` and `AddGrayscale()` by including them as JSON keys. [See sample usage](https://developer.pdf.co/api/pdf-to-image/png#ocrimagepreprocessingfilters). |
| `.AddGammaCorrection()` | array\[string (float format)] | `["1.4"]` | Adds a gamma correction filter to the image preprocessing pipeline used during OCR (Optical Character Recognition). This filter adjusts the brightness and contrast of an image by applying a non-linear gamma correction to improve text recognition quality. |
| `.AddGrayscale()` | boolean | `false` | Set to `true` to preprocessing filter that converts a colored document/image to grayscale before performing OCR. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# PDF Merging
Source: https://developer.pdf.co/integrations/n8n/merge-pdf
This operation merges multiple PDF files into a single PDF document. You can enable the auto-conversion feature to automatically convert DOC, DOCX, XLS, JPG, PNG, MSG, and EML files to PDF before merging.
The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed.
## Input
| Name | Description | Required |
| :-------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **Input Links** | A comma-separated list of files to merge. If you use a cloud service such as **Google Drive** or **Dropbox** ensure the links are publicly accessible. | Yes |
| **Automatically Convert Non-PDF Files** | Whether to automatically convert non-PDF files to PDF before the merging operation. Supported documents: **DOC, DOCX, XLS, JPG, PNG, MSG,** and **EML**. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :-------------------------------- | :------------- | :-------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `RenameMatchingFieldsDuringMerge` | boolean | `true` | This feature enables the renaming of field names during the merging of PDF files which contain forms. If set to `false`, it will retain the original field names. This is helpful for merged PDF forms with identical field names when the customer wants to auto-fill the identical field names in other pages. |
| `GenerateBookmarks` | boolean | `false` | This adds bookmarks to the merged document with names assigned to every merged document in the same order. |
| `BookmarkTitles` | array\[string] | - | An array containing the titles/names for bookmarks to be created. |
| `zipIncludeFilter` | string | - | You can control which files to include and exclude from input zip files with a profiles. |
| `zipExcludeFilter` | string | - | `zipIncludeFilter` and `zipExcludeFilter` support `*` and `?` wildcards. |
| `MergedDocumentTitle` | string | Title of the first document | Specifies a custom title for the merged document. Overrides the title of the first document during the merge process. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# PDF Security
Source: https://developer.pdf.co/integrations/n8n/pdf-add-remove-security
This operation allows you to add or remove security features such as passwords, accessibility restrictions, and modification limitations to your document.
**Modifying Restriction Settings**: To modify assembly or extraction settings that control PDF restrictions (such as `allowPrintDocument`, `allowFillForms`, `allowModifyDocument`, `allowAssemblyDocument`, and related permission settings), you **must** use the `ownerPassword` in your request. The `userPassword` alone cannot modify these permission settings. Attempting to change these restrictions with only a `userPassword` will result in an error: `"This file is password-protected. Please ensure you've entered the correct password…"`. This requirement applies only when making changes to those restrictions.
## Input
| Name | Description | Required |
| :--------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------------------------------------ |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Operation Mode** | Choose whether you want to add or remove security features from PDF documents. | Yes |
| **Owner password** | The main owner password that is used for document encryption and for setting/removing restrictions. | Fill at least one for Remove Security |
| **User Password** | The optional user password will be asked for viewing and printing document. | Fill at least one for Remove Security |
| **Encryption Algorithm** | Encryption algorithm. `AES_128bit` or higher is recommended. The available algorithms are: `RC4_40bit`, `RC4_128bit`, `AES_128bit`, `AES_256bit`. | No |
| **Allow Accessibility Support** | Allow content extraction for accessibility. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Document Assembly** | Allow document assembly. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Printing** | Allow printing. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Form Filling** | Allow form filling. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Document Modification** | Allow document modification. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Content Extraction** | Allow content extraction. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Allow Annotation Modification** | Allow annotation modification. Applies only if a User password is set, bypassed if the file is accessed using the Owner Password. | No |
| **Print Quality** | Specify the allowed printing quality. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Rotate PDF Pages
Source: https://developer.pdf.co/integrations/n8n/rotate-pdf
This operation allows you to rotate pages within your PDF file either manually or automatically by detecting and correcting rotation based on text analysis.
## Input
| Name | Description | Required |
| :--------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :-------------------------------------------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Rotation Mode** | Choose the rotation mode to apply. Use `Auto` to automatically detect and correct rotation based on text analysis, or manually specify the rotation angle. | Yes |
| **Rotation Angle** | Choose the rotation angle: `90°`, `180°`, or `270°`. | No, unless you’re using Manual Rotation mode. |
| **Pages** | Specify the pages you want to rotate using comma-separated values or page ranges (e.g., “`0,1,2-`” or “`1,2,3-7`”). Page indexing starts at `0`. Use “`!`” before a number to count from the end (e.g., “`!0`” for the last page). If left empty, all pages will be processed by default. The input must be a string. | No, unless you’re using Manual Rotation mode. |
| **OCR Language Name or ID** | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see Language Support. You can also use 2 languages simultaneously like this: `eng+deu` (any combination). | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [PDF.co Temporary Files Storage](https://developer.pdf.co/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Search and Replace/Delete Text
Source: https://developer.pdf.co/integrations/n8n/search-and-replace-text-in-pdf
This operation allows you to search for specific text and replace it with other text or images. It also enables you to search for specific text and delete it.
## Input
| Name | Description | Required |
| :--------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Operation Type** | Choose whether you want to search for text and replace it with other text or an image, or simply search and delete the matched text. | Yes |
| **Search Text** | Specify the text you wish to search for within the PDF document (e.g. company name). | Yes |
| **Replacement Text** | Enter the new text that will replace the found text. | Yes |
| **Replacement Image URL** | Provide the URL of the image to be used as a replacement for the located text. | Yes |
| **Pages** | Specify page indices as comma-separated values or ranges to process (e.g. “`0, 1, 2-`” or “`1, 2, 3-7`”). The first-page index is `0`. Use ”`!`” before a number for inverted page numbers (e.g. “`!0`” for the last page). If not specified, the default configuration processes all pages. The input must be in string format. | No |
| **Replacement Limit** | Limit the number of searches & replacements for every item. The value `0` means every found occurrence will be replaced. | No |
| **Use Regular Expressions** | Set to `true` to use regular expression for search string(s). | No |
| **Case Sensitive** | Set to `false` to don’t use case-sensitive search. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Links Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). | No |
| **Password** | The password of the password-protected PDF file | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{'YAdjustmentForReplacementText': '-1'}
```
With this input, the PDF.co operation will adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. Negative values for this parameter move text up, positive values move text down. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------------ | :------ | :------ | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AutoCropImages` | boolean | `false` | If you require to crop empty space around an inserted image use the following: `profiles": { 'AutoCropImages': true }` |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `YAdjustmentForReplacementText` | integer | - | Adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. See [**Adjust Text Alignment**](https://developer.pdf.co/api/pdf-search-text-and-replace/text#adjust-text-alignment) for more details. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Search in PDF
Source: https://developer.pdf.co/integrations/n8n/search-in-pdf
This operation enables you to search for specific text or keywords within your PDF documents or scanned images.
## Input
| Name | Description | Required |
| :----------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------- |
| PDF URL | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| Search Query | Specify the text you wish to search for within the PDF document (e.g. company name). | Yes |
| Use Regular Expressions | Set to true to enable regular expression search for the searchString(s) parameter. | No |
| Pages | Specify the page indices to search, using comma-separated values or ranges (e.g., “`0,1,2-`” or “`1,2,3-7`”). Page indexing starts at `0`. Use “!” before a number to count from the end (e.g., “`!0`” for the last page). Leave empty to search all pages. The input must be a string. | No |
| File Name | File name for the generated output, the input must be in string format. | No |
| Webhook URL | The callback URL or Webhook used to receive the output data. | No |
| Output Links Expiration (In Minutes) | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](https://developer.pdf.co/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| Inline | Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated. | No |
| Word Matching Mode | `WordMatchingMode` defines how search terms match PDF text. Modes: `None` (exact string match only), `SmartMatch` (default; flexible word boundary match, includes letters/digits/punctuation), `ExactMatch` (strict word boundaries, whole-word match only). | No |
| Password | The password of the password-protected PDF file | No |
| HTTP Username | HTTP auth user name if required to access source URL. | No |
| HTTP Password | HTTP auth password if required to access source URL. | No |
| Custom Profiles | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ 'OCRDetectPageRotation': true }
```
With this input, the PDF.co operation will rotate the scanned PDF automatically. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------------------------------- | :------ | :------------------------ | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ColumnDetectionMode` | string | `ContentGroupsAndBorders` | Controls column detection/alignment in PDF table extraction. Modes: `ContentGroupsAndBorders` (default; text + lines), `ContentGroups` (text grouping only), `Borders` (lines only), `BorderedTables` (OCR-based for bordered tables), `ContentGroupsAI` (AI for dense/complex layouts). |
| `DetectionMinNumberOfRows` | integer | 1 | Minimum number of rows to detect in a table |
| `DetectionMinNumberOfColumns` | integer | 1 | Minimum number of columns to detect in a table |
| `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | `0` | Maximum number of invalid subsequent rows allowed in a table |
| `DetectionMinNumberOfLineBreaksBetweenTables` | integer | `0` | Minimum number of line breaks between tables |
| `EnhanceTableBorders` | boolean | `true` | Enhance table borders or not |
| `OCRDetectPageRotation` | boolean | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See [**Support page rotation**](https://developer.pdf.co/api/pdf-find/basic#support-page-rotation) for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
You can also use `Custom Profiles` to:
* [Limit search to bordered tables only](https://developer.pdf.co/api/pdf-find/basic#find-only-bordered-tables)
* Visit this page for [general information on Profiles usage.](https://developer.pdf.co/api/profiles)
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# PDF Splitting
Source: https://developer.pdf.co/integrations/n8n/split-pdf
Split a single PDF file into multiple PDF files. This operation offers various methods for splitting, including by page number, page range, text search, or barcode search. This feature is particularly useful for segmenting large PDF documents or extracting specific sections.
## Input
| Name | Description | Required |
| :------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------------------------------------------- |
| **PDF URL to Split** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Split By** | Choose how you want to split your PDF file. You can split your PDF based on page numbers, search text, or barcode search. | Yes |
| **Page Number or Ranges** | Enter specific page numbers or ranges to extract. Use `1` for the first page, `1-3` for a range, or `7-` to include all pages from page 7 onward. Use negative numbers to count from the end (e.g., `-1` = last page, `-2` = second-to-last). Use `*` to split each page into a separate file. | No, unless you’re using Split by Pages |
| **Text Search String** | Enter the text to search for in PDF. Must be a String. | No, unless you’re using Split by Search Text |
| **Barcode Search String** | Enter the barcode macros string in PDF. See this guidance below to Split your PDF based on barcode search. | No, unless you’re using Split by Barcode |
| **Case-Sensitive Search** | Enable case-sensitive search. | No |
| **Regular Expression Search** | Enable regular expression search for the Search String parameter. | No |
| **Exclude Pages with Identified Text** | Exclude pages where the Search String text was found. | No |
| **OCR Language** | Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. | No |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **HTTP Username** | HTTP auth user name if required to access source URL. | No |
| **HTTP Password** | HTTP auth password if required to access source URL. | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](#custom-profiles) section to see all available parameters for your current endpoint. | No |
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Upload File
Source: https://developer.pdf.co/integrations/n8n/upload-file
This operation lets you upload your binary file, base64 data, or generated PDF to PDF.co and receive a URL in return. You can then use this URL with other PDF.co operations for further processing.
The Upload File operation acts as a bridge between the binary files or base64 data from the previous node and the PDF.co node, enabling smooth integration in automated workflows.
## Common Usage
These are typical workflows where files from popular sources are uploaded to PDF.co for further processing. Each example includes a step-by-step tutorial to help you set it up in n8n.
### Upload Binary File from Google Drive
You can get the file from Google Drive and upload it to PDF.co’s temporary storage before further processing.
### Upload Binary File from OneDrive
You can fetch files from OneDrive and send them to PDF.co’s temporary storage for use in other operations.
### Upload Binary File from IMAP
You can extract attachments from incoming emails and upload them to PDF.co’s temporary storage for processing.
## Input
| Name | Description | Required |
| :----------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- |
| **Upload Method** | Choose the upload method: use Standard Upload to upload files directly (supports up to 2GB), or `Base64` Upload to upload files using **base64-encoded** data. | Yes |
| **Binary File** | Enable this toggle when uploading binary file data from a previous node. | No |
| **Base64 Content** | The base64 encoded content of the file to upload. | Yes |
| **File Content** | Enter the text content for the PDF file to be created and uploaded to our storage. | Yes |
| **File Name** | File name for the uploaded PDF. The input must be a string. | Yes |
## Output
| Name | Description |
| :---- | :--------------------------------------------- |
| `url` | Direct URL to the final PDF file stored in S3. |
# URL/HTML to PDF Conversion
Source: https://developer.pdf.co/integrations/n8n/url-html-to-pdf
The URL/HTML to PDF Conversion operation lets you convert any publicly accessible webpage or raw HTML content into a downloadable PDF file. This is ideal for automating the creation of invoices, receipts, reports, or web page snapshots directly from HTML templates or live URLs. Simply provide either a valid URL or an HTML string as input. PDF.co will generate a high-quality PDF and return a file URL that you can use in the next step of your workflow—such as sending via email, storing in Google Drive, or combining with other documents.
## Input
| Name | Description | Required |
| :-------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------------------------------------------- |
| **Convert Type** | Choose the type of conversion you want to perform: `URL to PDF`, `HTML to PDF`, or `HTML Template to PDF`. If you’ve previously created an HTML template in PDF.co, you can select the template (by copying the Template ID) to inject it with dynamic JSON data and generate a PDF automatically. | Yes |
| **Source File URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/tools/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **HTML Code** | Enter the HTML code you want to convert into a PDF file. | No, unless you’re using HTML to PDF |
| **Template ID** | Set ID of HTML template to be used. View and manage your templates at HTML to PDF Templates. See guidance below to learn more about [**HTML Template to PDF**](#html-template-to-pdf) conversion. | No, unless you’re using HTML Template to PDF |
| **File Name** | File name for the generated output, the input must be in string format. | No |
| **Orientation** | Sets the document orientation. Options: `Portrait` for vertical layout, and `Landscape` for horizontal layout. | No |
| **Paper Size** | Specifies the paper size. Supports standard formats such as `Letter`, `Legal`, `Tabloid`, `Ledger`, and `A0` to `A6`. If you need a custom size, use the `Custom Paper Size `option instead. Do not use both `Paper Size` and `Custom Paper Size` at the same time. | No |
| **Custom Paper Size** | Set a custom size by providing width and height separated by a space, with optional units: `px` (pixels), `mm` (millimeters), `cm` (centimeters), or `in` (inches). Examples: ‘200 300’, ‘200px 300px’, ‘200mm 300mm’, ‘20cm 30cm’, ‘6in 8in’. | No |
| **Render Page Background** | Set to `false` to disable background colors and images are included when generating PDFs from HTML/URL | No |
| **Do Not Wait for Full Load** | Controls how thoroughly the converter waits for a page to load before converting HTML to PDF --- `false` waits for full page load, while `true` speeds up conversion by waiting only for minimal loading. | No |
| **Margins** | Set custom margins, overriding CSS default margins. Specify the margins in the format `{top} {right} {bottom} {left}`. You can use`px`,`mm`,`cm`or`in`units. Also, you can set margins for all sides at once using a single value. | No |
| **Media Type** | Controls how content is rendered when converting to PDF. Options: `print` (uses print styles), `screen` (uses screen styles), `none` (no media type applied). | No |
| **Header** | Add HTML code for the page header. This content will appear at the top of every page in the generated PDF. | No |
| **Footer** | Add HTML code for the page footer. This content will appear at the top of every page in the generated PDF. | No |
| **Webhook URL** | The callback URL or Webhook used to receive the output data. | No |
| **Output Link Expiration (In Minutes)** | Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from [**PDF.co Temporary Files Storage**](https://developer.pdf.co/api/file-upload/overview). The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using [**PDF.co Built-In Files Storage**](https://app.pdf.co/tools/files). | No |
| **Custom Profiles** | Use JSON to customize PDF processing with options like output resolution, OCR settings, text extraction methods, encryption, and image handling. Check our [Custom Profiles](https://app.pdf.co/tools/files) section to see all available parameters for your current endpoint. | No |
## HTML Template to PDF
HTML Template to PDF conversion allows you to transform predefined HTML templates into PDF documents with injected data. Each HTML template has a unique ID, you simply need to enter this ID in your n8n module and provide the required data in JSON format. You can create your own templates or use the predefined ones available in the [PDF.co HTML Templates Tool](https://app.pdf.co/html-templates-tool).
You can use data in JSON or CSV format to inject into your HTML template. See the sample below:
### JSON Sample
```
{'paid': true, 'invoice_id': '0002', 'total': '$999.99'}
```
### CSV Sample
```
`invoice_id, total\\r\\n12345,$999`
```
## Custom Profiles
You can set additional options for the operation used in the PDF.co node by using **Custom Profiles**. A custom profile is a string in **JSON-like** format containing predefined parameters.
Here’s an example of a Custom Profiles input:
```
{ "outputDataFormat": "base64" }
```
With this input, the PDF.co operation will return the output in `base64` format. You can find the list of available parameters for customizing profiles in the PDF.co operation documentation below:
You can use any regular API parameter from the [API Reference](/api) within n8n using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| :------------------------ | :----- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `HTMLCodeHeadInject` | string | - | Injects custom CSS and JavaScript code into the HTML `` section during conversion. See [HTML to PDF Knowledge Base](https://developer.pdf.co/knowledgebase/convert#pdf-from-html) for more information and sample usage. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See [**User-Controlled Encryption**](https://developer.pdf.co/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See [**User-Controlled Encryption**](/knowledgebase/user-controlled-encryption) for more information. |
## Output
| Name | Description |
| :-------------------- | :--------------------------------------------------------------------------------------------------------------------- |
| `jobId` | Unique identifier for the background job. |
| `pageCount` | Number of pages in the PDF document. |
| `error` | Indicates whether an error occurred (`false` means success) |
| `status` | Status code of the request (200, 404, 500, etc.). For more information, see [**Response Codes**](/api/response-codes). |
| `credits` | Number of credits consumed by the request |
| `remainingCredits` | Number of credits remaining in the account |
| `duration` | Time taken for the operation in milliseconds |
| `url` | Direct URL to the final PDF file stored in S3. |
| `name` | Name of the output file |
| `outputLinkValidTill` | Timestamp indicating when the output link will expire |
# Pabbly Connect
Source: https://developer.pdf.co/integrations/pabbly-connect
[Pabbly Connect](https://www.pabbly.com/connect/) is an online automation solution that connects two or more apps to manage and automate repetitive and manual tasks. It supports 800\+ apps integration for seamless real\-time data transfer.
## Pabbly Connect and PDF.co integration
## How to Fill a PDF Form
# Salesforce
Source: https://developer.pdf.co/integrations/salesforce
[Salesforce](http://salesforce.com/) provides customer relationship management services and also sells a complementary suite of enterprise applications focused on customer service, marketing automation, analytics, and application development.
## Salesforce and PDF.co integration
## Parse Invoice with Document Parser
# Sharepoint
Source: https://developer.pdf.co/integrations/sharepoint
[SharePoint](https://sharepoint.microsoft.com/) is a web\-based platform that integrates with other **Microsoft** services.
## SharePoint and PDF.co Integration
Please contact us to find out more:
## Extract Invoice Information
# Add Barcode
Source: https://developer.pdf.co/integrations/zapier/add-barcode-to-pdf
Enhance your Zapier workflow by integrating this step to generate a barcode and add it to an existing PDF document. You can also generate blank PDF files containing barcodes.
## Input
| Name | Description | Required |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Barcode Value** | Specify the value you wish to encode into the barcode. | Yes |
| **Barcode Type** | Select the [type of barcode](/api/barcode-reader) to generate. Defaults to a QR Code. | Yes |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. Leave this field empty to generate blank PDF files containing barcodes. | No |
| **X Co-ordinate** | Specify the X coordinate. Use the `PDF.co tool` to determine the `X` and `Y` coordinates on your PDF file. | Yes |
| **Y Co-ordinate** | Specify the Y coordinate. The [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper) can assist in finding the `X` and `Y` coordinates. | Yes |
| **Pages** | Indicate the pages where the barcode should be added using page numbers or ranges. Leave blank to include all pages. The first page is numbered `0`. Example: `0,2-5,7-`. | Yes |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{"OCRResolution": 600, "TrimSpaces": true, "OCRMode": "TextFromImagesAndFonts"}
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | -------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[`x`, `y`, `width`, `height`] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew, use the following: |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: `AES128`, `AES192`, `AES256`. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Add Form Field
Source: https://developer.pdf.co/integrations/zapier/add-formfield-to-pdf
Integrate this step into your Zapier workflow to add form fields to PDF documents, enhancing interactivity and data collection capabilities.
## Input
| Name | Description | Required |
| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Id** | Assign an ID to the form field. | Yes |
| **X Co-ordinate** | Specify the X coordinate for form field placement. Use the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper) to find `X` and `Y` coordinates. | Yes |
| **Y Co-ordinate** | Specify the Y coordinate. Assistance for coordinates can be found using the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper). | Yes |
| **Type** | Define the field type, such as TextField or Checkbox. | Yes |
| **Initial Value** | Set an initial value for the field. For checkboxes, use values like `X`, `true`, or `1` to mark them as checked. | No |
| **Size** | Define the field size, overriding initial settings. | No |
| **Pages** | Enter the page numbers (or ranges) from where the form field should be added. Leave empty to process all pages. The first page is `0`. Example: `0,2-5,7-`. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
# Add Image
Source: https://developer.pdf.co/integrations/zapier/add-image-to-pdf
Enhance your Zapier workflow by integrating this step to add images to PDF documents, offering flexibility and customization in document design.
## Input
| Name | Description | Required |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source File URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Image URL** | Specify the URL of the image to be added to the PDF. | Yes |
| **X Co-ordinate** | Determine the X coordinate for image placement. Use the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper) to find `X` and `Y` coordinates. | Yes |
| **Y Co-ordinate** | Specify the Y coordinate. Assistance for coordinates can be found using the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper). | Yes |
| **Pages** | Enter the page numbers (or ranges) from where the barcode should be added. Leave empty to process all pages. The first page is `0`. Example: `0,2-5,7-`. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
# Add Text
Source: https://developer.pdf.co/integrations/zapier/add-text-to-pdf
Integrate this step into your **Zapier** workflow to add text to **PDF** documents with a range of customization options.
## Input
| Name | Description | Required |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source File URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text to add** | Specify the text to be added. Use `\n` or `{{$$newLine}}` for line breaks, and leverage built-in macros like `{{$$PageNumber}}` and custom data macros. For using special macros on Zapier, see [Macros for Text](/knowledgebase/macros-for-text#special-macro-style%3A-square-brackets). | Yes |
| **X Co-ordinate** | Determine the `X` coordinate for text placement. Use the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper) to find `X` and `Y` coordinates. | Yes |
| **Y Co-ordinate** | Specify the `Y` coordinate. Assistance for coordinates can be found using the [PDF.co tool](https://app.pdf.co/pdf-edit-add-helper). | Yes |
| **Advanced Options** | Enable this to reveal additional fields | No |
| **FontName** | Select the font name from the [available font list](/knowledgebase/general). The default font is **Arial**. | No |
| **Pages** | Indicate specific page numbers or ranges where the text should be added. Leave blank to include all pages. The first page is numbered `0`. Example: `0,2-5,7-`. | No |
| **Link** | Enter a URL to make the text clickable. | No |
| **Font Size** | Set the font size, with the default being `12`. | No |
| **Color** | Choose the text color in hex format (`#RRGGBB` or `#AARRGGBB`, with `AA` as transparency). The default is `#000000`. | No |
| **Bold** | Enable this to apply bold styling to the font. | No |
| **Strikeout** | Enable this to apply a strikeout effect to the text. | No |
| **Underline** | Enable this to underline the text. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
# AI Invoice Parser
Source: https://developer.pdf.co/integrations/zapier/ai-invoice-parser
Enhance your Zapier workflow by integrating this step to automatically detect invoice and structurally extract data with our advanced AI. The AI Invoice parser automatically detects invoice layouts without the manual effort previously required to supply document parsing templates for reference.
**Important**
* **Only invoices will be parsed**. For all other documents, please use our existing [**Document Parser**](/integrations/zapier/document-parser).
* To ensure accurate processing, each invoice must be clearly separated. **If an invoice contains multiple pages, we recommend splitting it** into individual PDFs using the [Split PDF](/integrations/zapier/split-pdf) operation in your Zap.
* While AI Invoice Parser supports multi-page invoices, **the total page count for a single PDF must not exceed 100 pages**. Submitting large PDFs containing multiple invoices is not recommended.
## Input
| Name | Description | Required |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. Leave this field empty to generate blank PDF files containing barcodes. | Yes |
| **customField** | Comma-separated list of [custom field](/integrations/zapier/ai-invoice-parser#custom-fields) names to extract. Use `camelCase` for field names (e.g., `storeNumber`, `deliveryDate`). | No |
| **lineItemStructure** | A JSON object that defines a [custom structure](/integrations/zapier/ai-invoice-parser#line-item-structure) for line items. Each key is a field name (in `camelCase`) and each value is `"string"` or `"number"`. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
### Custom Fields
AI Invoice Parser with custom fields support automatically detects invoice layouts and extracts both standard schema data and user-specified custom fields without requiring manual templates.
The `customField` parameter allows you to specify additional fields to extract beyond the standard schema. Some examples include:
* `storeNumber` - Store or branch identifier
* `deliveryDate` - Expected delivery date
* `financialCharges` - Additional financial charges
* `lineTotal` - Total amount for line items
* `purchaseOrderRef` - Purchase order reference number
* `customerReference` - Customer reference number
* `departmentCode` - Department or cost center code
If a custom field returns an empty value, please [contact our support team](https://pdf.co/support/request?subject=ai-invoice-parser%20-%20custom%20fields) to help improve the extraction accuracy.
### Line Item Structure
The `lineItemStructure` parameter lets you define a custom schema for line items. Each key is a field name you choose (in `camelCase`) and each value is the expected data type — either `"string"` or `"number"`.
When provided, every object inside the `lineItems` array will contain exactly the fields you specified. If a value cannot be extracted from the invoice, the field will still be present with an empty or default value instead of being omitted.
Use `camelCase` for field names (e.g., `unitPrice`, `totalPrice`). The field names you define will be used as-is in the response, giving you full control over the output keys.
## Output
| Name | Description |
| ------------------ | ------------------------------------------------------------------------------ |
| `body` | An object array containing the all invoice parsing result. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
| `pageCount` | Total Page Count. |
## Supported Languages
* **Albanian (Shqip)**
* **Bosnian (Bosanski)**
* **Bulgarian (Български)**
* **Croatian (Hrvatski)**
* **Czech (Čeština)**
* **Danish (Dansk)**
* **Dutch (Nederlands)**
* **English**
* **Estonian (Eesti)**
* **Finnish (Suomi)**
* **French (Français)**
* **German (Deutsch)**
* **Greek (Ελληνικά)**
* **Hungarian (Magyar)**
* **Icelandic (Íslenska)**
* **Italian (Italiano)**
* **Latvian (Latviešu)**
* **Lithuanian (Lietuvių)**
* **Norwegian (Norsk)**
* **Polish (Polski)**
* **Portuguese (Português)**
* **Romanian (Română)**
* **Russian (Русский)**
* **Serbian (Српски)**
* **Slovak (Slovenčina)**
* **Slovenian (Slovenščina)**
* **Spanish (Español)**
* **Swedish (Svenska)**
* **Turkish (Türkçe)**
* **Ukrainian (Українська)**
# Anything to PDF
Source: https://developer.pdf.co/integrations/zapier/anything-to-pdf
Convert other file formats to **PDF**.
## Input
| Name | Description | Required |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input Format** | Choose the input for your format from the drop-down list. | Yes |
| **Input Source** | Provide the **URL** to the source document. If you use a cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output PDF Name** | The output file name. If left blank then the name `preview.pdf` will be used. | No |
| **Worksheet Index** | Used when converting from Spreadsheet. Enter the worksheet number you want to convert to **PDF**. The first worksheet is numbered `0`. If you want to process all worksheets, you can leave this field empty. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{"OCRResolution": 600, "TrimSpaces": true, "OCRMode": "TextFromImagesAndFonts"}
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Barcode Generator
Source: https://developer.pdf.co/integrations/zapier/barcode-generator
Add this step to your Zapier Workflow to generate a variety of barcodes, such as QR Codes.
## Input
| Name | Description | Required |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------- | -------- |
| **Barcode Value** | Specify the value to be encoded into the barcode. | Yes |
| **Barcode Type** | Choose the type of barcode to generate. Defaults to QR Code, with various other formats available. | No |
| **Generate Inline URL** | Opt for `true` to create an inline image URL, suitable for embedding in HTML or emails, and accessible offline. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{"OCRResolution": 600, "TrimSpaces": true, "OCRMode": "TextFromImagesAndFonts"}
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `Angle` | integer | `0` | See profiles.Angle |
| `NarrowBarWidth` | integer | `3` | See profiles.NarrowBarWidth |
| `CaptionFont` | string | `Arial, 12` | See profiles.CaptionFont |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Advanced Barcode Reader
Source: https://developer.pdf.co/integrations/zapier/barcode-reader
Integrate this step into your Zapier workflow to efficiently and accurately decode barcodes from images or PDF documents.
## Input
| Name | Description | Required |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input URL** | Provide the URL of the source file (PDF, PNG, JPG, TIFF) or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). Ensure the link is publicly accessible if using cloud services like Google Drive or Dropbox. | Yes |
| **Barcode Type to Read** | Select the barcode type for decoding. Defaults to QR Codes, with support for various other formats. | No |
| **Pages to Read From** | Specify page numbers or ranges for barcode reading. Leave blank to scan all pages. The first page starts at `0`. Example: `0,2-5,7-`. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------- |
| `barcode1` | An Object containing form barcode information such as `Barcode 1 Value`, `Barcode 1 Type Name`, `Barcode 1 Rect`, `Barcode 1 Page` etc. |
| `barcode2` | An object holding another barcode information, following the same pattern as `barcode1` for each output file. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{"OCRResolution": 600, "TrimSpaces": true, "OCRMode": "TextFromImagesAndFonts"}
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Compress & Optimize
Source: https://developer.pdf.co/integrations/zapier/compress
Add this step to your Zapier Workflow to compress & optimize your **PDF** files. This will reduce the file size of your **PDF** whilst maintaining the quality.
## Input
| Name | Description | Required |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the **URL** to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Output File Name** | The output file name. If left blank then the name `preview-compressed.pdf` will be used. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `JPEGQuality` | integer | `25` | Controls JPEG compression quality from 1 (worst quality, smallest size) to 100 (best quality, largest size). |
| `ImageOptimizationFormat` | string | - | The image optimization format. e.g. `"JPEG"` / `"Fax"` / `"Flate"` |
| `JPEGQuality` | integer | - | Quality setting for `"JPEG"` format. e.g. `0` - `100` |
| `ResampleImages` | boolean | - | Whether to resample the images or not. e.g. `true` / `false` |
| `ResamplingResolution` | integer | - | The DPI ([Dots-Per-Inch](https://en.wikipedia.org/wiki/Dots_per_inch)) setting for the document resampling. *e.g.* `72` / `96` / `120` / `150` / `200` |
| `GrayscaleImages` | boolean | - | Whether or not to remove color from images. e.g. `false` / `true` |
# Custom API Call
Source: https://developer.pdf.co/integrations/zapier/custom-api-call
Integrate this step into your Zapier workflow to execute custom API calls to the PDF.co API, designed for advanced users who need more flexibility and control.
## Input
| Name | Description | Required |
| -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **PDF.co API Endpoint** | Choose a PDF.co API endpoint or input a custom endpoint path. Consult the [PDF.co API Docs](/api) for detailed information. | Yes |
| **URL Input Parameter Override** | Override the `url` input parameter in the `Input JSON`. Accepts direct links, `filetoken://` links (for files in [PDF.co Built-In Files Storage](https://app.pdf.co/files)), or multiple comma-separated links. Compatible with Google Docs, Google Drive, Dropbox links that are accessible without login. | No |
| **Input JSON** | Provide an input JSON containing the necessary parameters for the API call. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary URL on the PDF.co file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
# Document Classifier
Source: https://developer.pdf.co/integrations/zapier/document-classifier
Integrate this step into your Zapier workflow to analyze the text of a document using AI and classify it into categories like invoice, order, or industry. This feature is useful for quickly identifying the origin of a document and can be customized with specific rules.
## Input
| Name | Description | Required |
| -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Input Document URL** | Provide the URL of the input PDF document or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For other cloud services like Google Drive or Dropbox, ensure the link is publicly accessible. | Yes |
| **Custom Classification Rules** | Define classification rules in CSV format, if required. Format per row: `classname,logic,keyword1,keyword2`. Example: `Amazon,AND,Amazon AWS,AWS Invoice`. Refer to [PDF Classifier](https://pdf.co/pdf-classifier) for detailed instructions. | No |
| **CSV Rules URL** | Provide a URL to a CSV file containing custom classification rules. The format for each row should be: `classname,logic,keyword1,keyword2`. Example: `Amazon,AND,Amazon AWS,AWS Invoice`. | No |
| **Enable Case Sensitive Custom Rules** | Indicate if the keywords in the custom rules should be case sensitive. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary URL on the PDF.co file server. |
| `classes` | An array containing possible category of input document. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------ | ----------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RenderTextObjects` | boolean | `true` | Render text objects or not |
| `RenderVectorObjects` | boolean | `true` | Render vector objects or not |
| `RenderImageObjects` | boolean | `true` | Render image objects or not |
| `TIFFCompression` | string | `LZW` | TIFF compression algorithm. The options are: None, LZW, CCITT3, CCITT4, RLE |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. |
| `RotationAngle` | integer | `-` | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
# Document Parser
Source: https://developer.pdf.co/integrations/zapier/document-parser
Integrate this step into your Zapier workflow to extract data from invoices, reports, statements, and other documents using AI-based templates. This is ideal for automating data extraction tasks, significantly reducing manual workload.
## Input
| Name | Description | Required |
| ------------------- | ---------------------------------------------------------------------------------------------------------------- | -------- |
| **Document URL** | Provide the URL of the PDF document to be processed. | Yes |
| **Template Id** | Specify your PDF.co [template IDs](https://app.pdf.co/document-parser/templates) to use for parsing. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | --------------------------------------------------------------------------------------------------- |
| `body` | An object array containing the all document parsing result. |
| `simplifiedData` | An object containing more simplified and compact result. It is easy to consume in next Zapier step. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Send Email With Attachment
Source: https://developer.pdf.co/integrations/zapier/email-send
Integrate this step into your Zapier workflow to send emails with attachments, providing a seamless way to share documents and information.
## Input
| Name | Description | Required |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------- | -------- |
| **Sender Email** | Specify the email address of the sender. | Yes |
| **Recipient Email** | Enter the email address of the recipient. | Yes |
| **Email Subject** | Define the subject of the email. | Yes |
| **SMTP Server** | Provide the address of the SMTP server. | Yes |
| **SMTP Server Port** | Indicate the port number used by your SMTP server. | Yes |
| **SMTP Username** | Enter your username for the SMTP server. | Yes |
| **SMTP Password** | Provide your password for the SMTP server. | Yes |
| **CC** | Add Carbon Copy (CC) recipients if required. | No |
| **BCC** | Add Blind Carbon Copy (BCC) recipients if necessary. | No |
| **Email Body (Plain Text)** | Input the content of the email in plain text. | No |
| **Email Body (HTML)** | Enter the email content in HTML format. | No |
| **Attachment URLs** | Provide a comma-separated list of URLs for any attachments to be included in the email. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
## Output
| Name | Description |
| ------------------ | ------------------------------------------------------------------------------ |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# PDF Find Table
Source: https://developer.pdf.co/integrations/zapier/find-table
Integrate this step into your Zapier workflow to utilize AI technology in scanning PDF documents for tables. It returns an array of found tables, including coordinates and details about the detected columns.
## Input
| Name | Description | Required |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- |
| **PDF Document URL** | Enter the URL of the input PDF document or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If using other cloud services like Google Drive or Dropbox, ensure the link is publicly accessible. | Yes |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| `url` | The temporary URL on the PDF.co file server. |
| `table` | An object array containing the found table results. |
| `TableData.table_1` | An object that includes a found table result. |
| `TableData.table_2` | An object holding another search result for table, following the same pattern as `TableData.table_1` for each output file. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| -------------------------------------------------- | ------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ColumnDetectionMode` | string | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. Refer to the [Column Detection Mode](#column-detection-mode) section for more information. |
| `DetectionMinNumberOfRows` | integer | `1` | Minimum number of rows to detect in a table |
| `DetectionMinNumberOfColumns` | integer | `1` | Minimum number of columns to detect in a table |
| `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | `0` | Maximum number of invalid subsequent rows allowed in a table |
| `DetectionMinNumberOfLineBreaksBetweenTables` | integer | `0` | Minimum number of line breaks between tables |
| `EnhanceTableBorders` | boolean | `true` | Enhance table borders or not |
| `OCRDetectPageRotation` | boolean | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See Support page rotation for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `requestParametersDocument` | string | - | The document to be processed. |
| `responseParameters` | object | - | The parameters returned in the response. |
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
# Getting Started with Zapier
Source: https://developer.pdf.co/integrations/zapier/getting-started
Connect **PDF.co** to thousands of other apps with **Zapier**.
[Zapier](https://zapier.com/apps/pdfco/integrations) is the no-code automation tool that lets you connect **PDF.co** to 2,000+ other web services. Automated connections called Zaps, set up in minutes with no coding, can automate your day-to-day tasks and build workflows between apps that otherwise wouldn’t be possible.
Each Zap has one app as the Trigger, where your information comes from and which causes one or more Actions in other apps, where your data gets sent automatically.
Sign up for a free [Zapier](https://zapier.com/apps/pdfco/integrations) account, and from there you can jump right in. To help you hit the ground running, here are some popular pre-made Zaps.
## How to Connect PDFco to Zapier
* Log in to your Zapier account or create a new account.
* Navigate to “My Apps” from the top menu bar.
* Now click on “Connect a new account…” and search for “PDF.co”
* Use your credentials to connect your PDF.co account to Zapier.
* Once that’s done you can start creating an automation! Use a pre-made Zap or create your own with the Zap Editor. Creating a Zap requires no coding knowledge and you’ll be walked step-by-step through the setup.
* Need inspiration? See everything that’s possible with [Zapier and PDF.co Integrations](https://zapier.com/apps/pdfco/integrations).
## Convert PDF Files to CSV with Zapier and PDF.co
# HTML to PDF
Source: https://developer.pdf.co/integrations/zapier/html-to-pdf
Convert HTML to PDF. Also supports HTML templates.
## Input
| Name | Description | Required |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------ |
| **HTML or URL Input** | Enter the **HTML** content or **URL** which you want to convert. | Yes (unless **HTML Template ID** is used) |
| **HTML Template ID** | Select a Template ID from your [HTML to PDF Templates](https://app.pdf.co/templates/html) | Yes (unless **HTML or URL Input** is used) |
| **Page Orientation** | Choose the **PDF** page orientation. | No |
| **Page Size** | Select the paper size for the **PDF**. | No |
| **Custom Page Size** | Use this instead of **Page Size** to specify a custom paper size for the **PDF** in `width height` format. You can use `px`, `mm`, `cm` or `in` units. For example: `200px 300px`, `200mm 300mm`, `20cm 30cm`, or `6in 8in`. | No |
| **Custom Margins** | Override the default margins. Specify the margins in the `top right bottom left` order. You can use `px`, `mm`, `cm` or `in` units. You can set margins for all sides at once using a single value, for example: `10px`. | No |
| **Render Page Background** | Whether to render the page background from the **HTML** source or not. | No |
| **Media Type** | Select `Print` or `Screen` quality for the **PDF** output | No |
| **Do not wait until full page load** | Set to `true` to not wait for full page load. Helps speed up pages with dynamic content. | No |
| **HTML Template Data** | Input data for **HTML** templates. Accepts **JSON** or **CSV** format. For example, for **JSON**: `'{ "total": "500" }'`, or for **CSV**: `'column1,column2,column3 value1,value2,value3'`. If your template ID has variables this is how you can input data your into the **PDF**. | No |
| **Name** | The output file name. If left blank then `htmltopdf.pdf` will be used. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `HTMLCodeHeadInject` | string | - | Injects CSS into the HTML `` section to prevent page breaks within specified elements. See HTML to PDF Knowledge Base for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Integrating File Sources with PDF.co
Source: https://developer.pdf.co/integrations/zapier/input-file-sources
Incorporating file sources such as Google Drive, Dropbox, OneDrive, and Box into your Zapier workflows is essential when working with PDF.co. This guide details the specific input properties from each of these services to be used with PDF.co’s Zapier plugin.
## URL Accessibility for PDF.co
The API supports TLS 1.2 and 1.3 for secure connections. Earlier versions such as TLS 1.0 and 1.1 are deprecated and should be avoided.
For successful integration, any publicly accessible URL can be used. Whether you’re using services like Google Drive, Dropbox, or others, make sure the URLs are accessible to PDF.co. We recommend using Zapier-generated URLs (detailed in the sections below) as they are guaranteed to be compatible with PDF.co. If you’re using direct URLs not provided via Zapier, please verify their public accessibility to ensure PDF.co can process your files efficiently.
## Google Drive
When integrating Google Drive with PDF.co in Zapier, use the File property to retrieve files. This property allows PDF.co to access the file needed for processing.
## Dropbox
**For Dropbox, the `Direct Media Link` property is utilized. However, it’s important to be aware of Dropbox’s file size limitation:**
* Limitation: Access to files is restricted up to 100 MB only. Larger files may result in an error.
## OneDrive
OneDrive integration requires the `Download URL` property. This is effective for accessing files and passing them to PDF.co.
## Box
Box, similar to Google Drive, uses the `File` property to share files with PDF.co. This ensures seamless file transfer within your Zapier workflow.
## SharePoint
For SharePoint, use the `Direct Download Link` property. Ensure the link is publicly accessible.
**Steps to Generate the Direct Download Link:**
**Add the “Microsoft SharePoint” Trigger**
* In SharePoint, add the **“New file in folder”** Trigger Event to create a shareable link that can be adjusted for direct access.
* **Site Address**: Select your SharePoint site.
* **Library Name**: Choose the library containing your file.
* **Item Id**: Specify the file you want to share.
For more information, see [Microsoft Docs on creating sharing links](https://docs.microsoft.com/en-us/connectors/sharepointonline/#create-sharing-link-for-a-file-or-folder).
If you are not using a Zapier trigger/action or encountering any issues, you can manually modify the link for access compatibility:
1. **Modify the Link for PDF.co Compatibility**
* Append `?download=1` to convert it into a direct download link compatible with PDF.co.
**Example:**
Original Link:
```
https://yourcompany.sharepoint.com/:b:/s/SharedDocuments/Ed1...4ig
```
Modified Link:
```
https://yourcompany.sharepoint.com/:b:/s/SharedDocuments/Ed1...4ig?download=1
```
2. **Troubleshooting**:
* Ensure the sharing settings allow public access.
* Test the link in a browser to confirm it downloads the file directly.
* Double-check permissions if PDF.co cannot access the file.
## PDF.co File Storage
# Merge PDF
Source: https://developer.pdf.co/integrations/zapier/merge
Add this step to your Zapier Workflow to combine multiple document formats into a single **PDF**. Supported formats include **PDF, DOC, DOCX, XLS, JPG, PNG** and more.
The total combined size of all input file URls must not exceed **2 GB**. Requests that exceed this limit will not be processed.
## Input
| Name | Description | Required |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- |
| **Source Files URLs** | A comma-separated list of files to merge. If you use a cloud service such as **Google Drive** or **Dropbox** ensure the links are publicly accessible. | Yes |
| **Automatically Convert Non-PDF Files** | Whether to automatically convert non-PDF files to **PDF** before the merging operation. | No |
| **Output File Name** | The output file name. If left blank then the name of the last file in the **Source PDF URL** list will be used. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| --------------------------------- | -------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `RenameMatchingFieldsDuringMerge` | boolean | true | This feature enables the renaming of field names during the merging of PDF files which contain forms. If set to false, it will retain the original field names. This is helpful for merged PDF forms with identical field names when the customer wants to auto-fill the identical field names in other pages. |
| `GenerateBookmarks` | boolean | `false` | This adds bookmarks to the merged document with names assigned to every merged document in the same order: |
| `BookmarkTitles` | array\[string] | - | An array containing the titles/names for bookmarks to be created |
| `zipIncludeFilter` | string | - | You can control which files to include and exclude from input zip files with a profiles. |
| `zipExcludeFilter` | string | - | zipIncludeFilter and zipExcludeFilter support \* and ? wildcards. |
| `MergedDocumentTitle` | string | Title of the first document | Specifies a custom title for the merged document. Overrides the title of the first document during the merge process. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# PDF Filler
Source: https://developer.pdf.co/integrations/zapier/pdf-filler
Integrate this step into your **Zapier** workflow to:
* Add text, images and signatures to a **PDF**
* Fill **PDF** form fields
* Create a new **PDF** from a template
## Input
| Name | Description | Required |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------- |
| **Source PDF** | Provide the URL of the PDF document, or leave blank to create a new PDF from scratch. | No |
| [**Text Annotations**](#text-annotations) | Specify each text annotation using the format `x;y;page;text`. See [Text Annotations](#text-annotations) for more information. | Fill at least one |
| [**Image Embeds**](#image-embeds) | Define each image using the format `x;y;page;urltoimage;link;width;height`. Use PDF.co [PDF Inspector](https://app.pdf.co/pdf-edit-add-helper) to find or measure PDF coordinates. | Fill at least one |
| [**Fillable Form Fields**](#fillable-form-fields) | Specify each fillable field value in the format `page;fieldName;value`. Use the [PDF.co Info Tool](https://app.pdf.co/pdf-info) or the **Get PDF Info** PDF.co step for field names in PDF forms. | Fill at least one |
| **Output PDF Name** | Name of the output PDF file. | No |
| **Allow Empty Text and Image Objects** | Enable to process text objects and images even with empty URLs. | No |
| **Auto-trim Input Values** | Automatically removes leading/trailing spaces and line breaks from input text values. | No |
| **Template Data** | Input data in `JSON` format that can be used within annotations and fields. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
## Data Format and Examples
### Text Annotations
This parameter represents one or more text objects to add to a **PDF**. Each text object is made of a parameter separated by the ; symbol. It uses the format:
`x;y;pages;text;fontSize;fontName;fontColor;link;transparent;width;height;alignment`
#### **EXAMPLE** Sample code
Adds a text annotation, which links to `www.pdf.co`, to the `20,20` coordinate of all pages in a document (by using: `0-`). The text annotation uses the `Arial` font in red (by using `FF0000`) with a font size of `24`. It has a transparent background, a defined `300` by `200` bounding box and with text aligned to the `right`.
```text theme={null}
20;20;0-;Test Text;24;Arial;FF0000;www.pdf.co;true;300;200;right
```
#### **EXAMPLE** Sample code
Where `24` is the font size. You can also add styles along with the font size using the following modifiers:
* `+bold`
* `+italic`
* `+underline`
* `+strikeout`
```text theme={null}
20;20;0-;Testing Text;24+bold+italic;Arial
```
#### **EXAMPLE** Sample code
Another example with `bold`, `italic`, `underline` and `strikeout` styles would be as follows:
```text theme={null}
250;20;0-;Text1|250;30;0-;Text2|250;50;0-;Text3
```
#### **EXAMPLE** Sample code
To put multiple objects, just use the `|` separator between objects.
```text theme={null}
20;20;0-;Test Text;24;Arial;FF0000;www.pdf.co;true;300;200;right
```
If you need to insert a line break then use `\n` or `{{$$newLine}}`.
### Image Embeds
Each image or PDF object can be defined as:
`x;y;pages;urlToImageOrPDF;linkToOpen;width;height`
#### **EXAMPLE** Sample code
```text theme={null}
20;80;0-;bytescout-com.s3-us-west-2.amazonaws.com/files/pdf-edit/logo.png;www.pdf.co;200;200
```
#### **EXAMPLE** Sample code
To put multiple objects, just use the `|` separator between objects.
```text theme={null}
100;180;0-;bytescout-com.s3-us-west-2.amazonaws.com/files/pdf-edit/logo.png|400;180;0-;bytescout-com.s3-us-west-2.amazonaws.com/files/pdf-edit/logo.png;www.pdf.co;200;200
```
You can also use a [base64 datauri](https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/Data_URLs) embedded image or a `filetoken://` link to a file from [PDF.co Built-In Files Storage](https://app.pdf.co/files).
### Fillable Form Fields
To fill fields in a **PDF** form, use the format `page;fieldName;value`.
To define the `font name`, `size` and `style` for the form input, use the following schema:
`page;fieldName;Field Text;size+bold+italic+underline+strikeout;FontName`
#### **EXAMPLE** Sample code
```text theme={null}
0;editbox1;text for my edit box;12+bold;Arial
```
#### **EXAMPLE** Sample code
To fill a checkbox, use `true` against the target checkbox object.
```text theme={null}
0;checkbox1;true
```
#### **EXAMPLE** Sample code
To separate multiple objects, use the `|` separator.
```text theme={null}
0;editbox1;text for my edit box;12+bold;Arial|0;editbox2;text for another edit box;12;Arial|0;checkbox1;true
```
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | -------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `Pages[0].SetCropBox()` | array\[string] | - | Crop a PDF file using an array to define the crop area. The crop box is defined by a rectangle \[x, y, width, height] in PDF points (1 Point = 1/72 inches). |
| `DisableLigatures` | boolean | `false` | To disable ligaturization, for example for Hebrew, use the following: |
| `FlattenDocument()` | boolean | `false` | Flattening a document renders it as read-only. Handy if you want to remove editing or copying capability. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Get PDF Information
Source: https://developer.pdf.co/integrations/zapier/pdf-info
Integrate this step into your Zapier workflow to extract information from a PDF document, including form fields, page count, size, author, description, keywords, and more.
## Input
| Name | Description | Required |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **PDF URL** | Provide the URL of the source PDF document or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like Google Drive or Dropbox, ensure the link is publicly accessible without a password. | Yes |
| **Extract Fillable Fields** | Enable this option (`true`) to extract information about fillable fields in the PDF form. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `info` | An object containing the PDF information such as `Title`, `PageCount`, `Author`, `Subject`, `CreationDate`, etc. It also contains the `FormField` object array in case we have the `Extract Fillable Fields` property enabled. |
| `SimplifiedFieldsData.field_1` | An Object containing form field information such as `FieldName`, `Value`, `Type`, `AltFieldName` and `PageIndex`. Please note: this output object will only produced if input is configured for `Extract Fillable Fields`. |
| `SimplifiedFieldsData.field_2` | An object holding another form field information, following the same pattern as `SimplifiedFieldsData.field_1` for each output file. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------ | ----------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
# PDF Page Tools
Source: https://developer.pdf.co/integrations/zapier/pdf-page-tools
Enhance your Zapier workflow by integrating this step for functionalities like rotating and deleting pages from a PDF document.
## Input
| Name | Description | Required |
| -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **PDF Source Link** | Provide a URL to the PDF file or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like Google Drive or Dropbox, ensure the link is publicly accessible. | Yes |
| **Mode** | Select the operation mode: rotate pages or delete pages. | No |
| **Rotation Angle** | Define the rotation angle in degrees for the `Rotate Pages` mode. AI-based auto-rotation based on text analysis is also available. | No |
| **OCR Language for auto-rotate** | Choose the OCR language for text recognition in scanned PDFs, primarily for auto-rotation. Default is English. | No |
| **Page Numbers** | Indicate the pages to process. Use comma-separated page numbers or ranges. Note: For `Rotate Pages` mode, the first page is `0`; for `Delete Pages`, it starts from `1`. Example: `1,2-5,7-`. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# PDF Security
Source: https://developer.pdf.co/integrations/zapier/pdf-password-and-security
Integrate this step into your Zapier workflow to add or remove password protection and security features to or from a PDF file.
**Modifying Restriction Settings**: To modify assembly or extraction settings that control PDF restrictions (such as `allowPrintDocument`, `allowFillForms`, `allowModifyDocument`, `allowAssemblyDocument`, and related permission settings), you **must** use the `ownerPassword` in your request. The `userPassword` alone cannot modify these permission settings. Attempting to change these restrictions with only a `userPassword` will result in an error: `"This file is password-protected. Please ensure you've entered the correct password…"`. This requirement applies only when making changes to those restrictions.
## Input
| Name | Description | Required |
| -------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| **Source PDF URL** | Provide the URL of the source PDF document or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like Google Drive or Dropbox, ensure the link is publicly accessible. | Yes |
| **Operation Mode** | Select to either add or remove security features from the PDF file. | Yes |
| **Owner Password** | Set the owner password for applying restrictions and encryption. You can also utilize [user-controlled data encryption](/knowledgebase/user-controlled-encryption). | Fill at least one for Remove Security |
| **User Password** | Optionally set a user password required to view or print the document. | Fill at least one for Remove Security |
| **Encryption Algorithm** | Select the encryption algorithm for built-in PDF encryption. AES-128 or higher is recommended. | No |
| **Allow Printing** | Determine if the PDF can be printed. Applies when the user password is entered; bypassed with the owner password. | No |
| **Define Printing Quality** | Specify the allowed printing quality. Applies with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Document Assembly** | Decide if document assembly is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Content Copying** | Decide if content copying is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Content Copying for Accessibility** | Decide if content extraction for accessibility is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Document Modification** | Decide if document modification is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Form Field Filling** | Decide if filling form fields is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Enable Commenting** | Decide if interacting with text annotations and forms is allowed. Effective with `userPassword`; bypassed with `ownerPassword`. | No |
| **Specify Name** | Provide a base file name for the processed PDF files. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Convert Scanned PDF to Searchable PDF
Source: https://developer.pdf.co/integrations/zapier/pdf-searchable
Enhance your Zapier workflow by integrating this step to convert scanned PDFs or image files into text-searchable PDF documents, leveraging advanced OCR technology.
## Input
| Name | Description | Required |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source File URL** | Provide the URL of the source file in PDF, PNG, TIFF, or JPG format. Also accepts a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). For cloud services like Google Drive or Dropbox, ensure the link is publicly accessible. | Yes |
| **OCR Language** | Specify the OCR language for text extraction from scanned documents. Default is English. | No |
| **Output File Name** | Name for the converted searchable PDF. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------ | ----------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. |
| `OCRResolution` | integer | `300` | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
# PDF to Anything
Source: https://developer.pdf.co/integrations/zapier/pdf-to-anything
Convert PDF to other file formats.
## Input
| Name | Description | Required |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Output Format** | Select the format you want to convert your **PDF** to from the drop down list in the workflow. | Yes |
| **Source PDF URL** | Provide the **URL** to the source **PDF** document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Page Selection** | Specify the pages or ranges you want to process. Enter a comma-separated list (e.g., `0,1-2,5-`). Leave blank to include all pages. Note: Page indexing starts at `0`. | No |
| **Output File Name** | The output file name. If left blank then the name of the last file in the **Source PDF URL** list will be used. | No |
| **Inline Output Option** | Set to `true` to receive the extracted content directly as a body variable. By default, a link to the output file will be returned in the url object in the return `JSON`. | No |
| **OCR Language for Scanned Documents** | Choose the OCR (Optical Character Recognition) language for extracting text from scanned **PDF**, **PNG**, **JPG** documents. The default language is English. | No |
| **Extraction Region** | Define coordinates for extraction with a list of comma-separated `x`, `y` coordinate numbers, e.g, `0, 0, 100, 100` is a square 100 pixels in from the top left corner of a **PDF** document. | No |
| **Enable Line Grouping** | Set to `true` to enable line grouping within table cells. | No |
| **Unwrap** | Set to `true` to unwrap lines into single line within table cells. Works only when **Enable Line Grouping** is enabled. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description | Available for |
| ------------------------------ | ----------------------------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `ColumnDetectionMode` | string | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. See [Column Detection Mode](#column-detection-mode) for more information. | PDF to CSV, PDF to XLS |
| `OCRMode` | string | `Auto` | Specifies how OCR (Optical Character Recognition) should process input content, offering various modes to tailor text extraction based on content type such as images, fonts, and vector graphics. For more information, see OCR Extraction Modes. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OCRResolution` | integer | 300 | Use this parameter to change the OCR resolution from the default 300 dpi. The range is from 72 to 1200 dpi. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RotationAngle` | integer | - | Use manual rotation to handle PDFs with vertically drawn text. Normally, OCR automatically detects page rotation in PDFs and extracts text accurately. However, in some cases, the PDF might not have an actual rotated page --- Rather, the text itself is drawn vertically. In such scenarios, auto-detection may fail. You can use this parameter to manually set the page rotation. The available angles are: 0, 1, 2, 3. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `LineGroupingMode` | string | `None` | Controls line grouping in PDF text extraction. Modes: None (no grouping), GroupByRows (merge rows if all cells align), GroupByColumns (merge cells by column), JoinOrphanedRows (merge single-cell rows to above if no separator). | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `ConsiderFontColors` | boolean | `false` | Controls whether font colors should be considered when detecting table structure and merging text objects during PDF extraction. Set to true to consider font colors. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DetectNewColumnBySpacesRatio` | string | `1.2` | Controls how spaces between words are interpreted for column detection in PDF text extraction. It defines the ratio of space width that determines when text should be treated as being in separate columns. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `AutoAlignColumnsToHeader` | boolean | `true` | Controls how columns are detected and aligned during table extraction from PDF documents. It affects both table structure detection and text extraction with formatting preservation. Set to true to automatically align columns to the header row. When set to true (default), the row with the most columns is used as the header, and all other rows are aligned to this structure --- ideal for well-structured tables. When set to false, columns are analyzed independently across all rows to build the structure, which works better for inconsistent or irregular tables. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OCRImagePreprocessingFilters` | object | - | Image preprocessing filters for OCR. Refer to [OCRImagePreprocessingFilters](#ocrimagepreprocessingfilters) for usage examples. | |
| `.AddGrayscale` | boolean | `false` | Converts to grayscale before OCR. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `.AddGammaCorrection` | array\[string (float format)] | \["1.4"] | Adds a gamma correction filter. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RenderTextObjects` | boolean | `true` | Controls whether to render text objects in the PDF document. When set to true, it will render all text objects in the PDF document. Set to false to skip over text objects during rendering. See Disable Text Layer for more information. | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `RenderImageObjects` | boolean | `true` | Render image objects or not | PDF to JPG, PDF to PNG, PDF to WEBP |
| `RenderVectorObjects` | boolean | `true` | Render vector objects or not | PDF to JPG, PDF to PNG, PDF to WEBP |
| `JPEGQuality` | integer | `85` | See profiles.JPEGQuality | PDF to JPG |
| `WEBPQuality` | integer | `75` | See profiles.WEBPQuality | PDF to WEBP |
| `TIFFCompression` | string | `LZW` | See profiles.TIFFCompression | PDF to TIFF |
| `RenderingResolution` | integer | `120` | See Set Image Resolution for more information. | PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `OptimizeImages` | boolean | `true` | Some PDF may have high quality images used in the document and you may need to keep the quality of these images in the output HTML. By default PDF to HTML is optimizing images and you can easily turn it off. See Control Image Quality for more information. | PDF to HTML |
| `OutputPageWidth` | integer | `1024` | Control page width (in pixels) for output HTML. Height is calculated and used according to the original pdf pages ratio. See Control Output Page Width for more information. | PDF to HTML |
| `AdditionalCssStyles` | string | `“` | To inject CSS for layout options in your HTML. Example: `#canvas { zoom: 50%; }`. Scale the div that contains all generated HTML pages by 50%. See Inject CSS for more information. | PDF to HTML |
| `SaveVectors` | boolean | `false` | Controls whether to save vector graphics during PDF to HTML conversion. Set to true to save vector graphics. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `SaveImages` | string | `None` | Controls how images are saved during PDF to HTML conversion. Modes: None (no images), OuterFile (save to sub-folder), Embed (embed as Base64 data:URI). | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML, PDF to HTML |
| `ConsiderFontSizes` | boolean | `false` | Set to true to this parameter makes the converter consider font size differences in document text when detecting and parsing table structures. This can be helpful in cases where tables are formatted using different font sizes to distinguish between headers, data cells, or other structural elements. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `ExtractionArea` | array\[number] | - | Extract text in a specific area by defining the extraction area - set with points in the format \[x, y, width, height]. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `ExtractShadowLikeText` | boolean | `true` | Controls whether to extract invisible text from a PDF document. Set to false to skip over invisible text during extraction. This is particularly useful when dealing with PDFs that contain hidden text layers or when you only want to extract visible content. When this value is set to false, OCRMode must be set to Auto to properly apply the shadow text filtering effect. | PDF to CSV, PDF to JSON, PDF to XLS, PDF to XML |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. | PDF to CSV, PDF to JSON, PDF to Text, PDF to XLS, PDF to XML, PDF to HTML, PDF to JPG, PDF to PNG, PDF to WEBP, PDF to TIFF |
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
### `OCRImagePreprocessingFilters`
To set image preprocessing filters, please use:
```json theme={null}
{
"profiles": "{
"ExtractShadowLikeText": false,
"OCRMode": "Auto",
"OCRImagePreprocessingFilters.AddGrayscale()": [],
"OCRImagePreprocessingFilters.AddGammaCorrection()": [
1.4
]
}"
}
```
# Search and Delete Text
Source: https://developer.pdf.co/integrations/zapier/search-and-delete-text
Enhance your Zapier workflow by integrating this step to locate and eliminate specified text from a PDF document. This feature supports the use of regular expressions for advanced text matching.
## Input
| Name | Description | Required |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text to Search and Remove** | Define the text to be searched for and removed within the PDF. A maximum of five search-and-delete operations are permitted per task. | Yes |
| **Enable Regular Expressions** | Activate this option to use regular expressions in your search. For example, `[0-9]{3}-[0-9]{2}-[0-9]{4}` can be used to locate SSN patterns. | No |
| **Case-Sensitive Search** | Toggle this option to make the search case-sensitive. | No |
| **Pages** | Specify the pages to be searched. Enter a comma-separated list of page numbers or ranges. Note: The first page is numbered `0`. For example: `0,1-5,7-`. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `removeTextUnderPatch` | boolean | `true` | Controls whether to remove text under the patch or not |
| `usepatch` | boolean | `false` | Controls whether to use a patch or not |
| `patchColor` | string | `#000000` | Controls the color of the patch |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Search and Replace Text
Source: https://developer.pdf.co/integrations/zapier/search-and-replace-text
Enhance your Zapier workflow by integrating this step to locate and replace specific text within a PDF document. This feature also supports the use of regular expressions for advanced text manipulation.
## Input
| Name | Description | Required |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text to Locate** | Define the text to be searched for within the PDF. A maximum of five search-and-replace operations per task are allowed. | Yes |
| **Replacement Text** | Specify the text that will replace the located text. Ensure the number of search and replacement text entries correspond. The operations will be executed in the order they are entered. | Yes |
| **Enable Regular Expressions** | Activate this option to use regular expressions in your search, such as `[0-9]{3}-[0-9]{2}-[0-9]{4}` to find SSN patterns. | No |
| **Case-Sensitive Search** | Toggle this for a case-sensitive search. | No |
| **Target Pages** | Specify the pages to be searched. Enter a comma-separated list of page numbers or ranges. Note: The first page is numbered `0`. For example: `0,1-5,7-`. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AutoCropImages` | boolean | `false` | If you require to crop empty space around an inserted image use the following: `profiles": { 'AutoCropImages': true }` |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `YAdjustmentForReplacementText` | integer | - | Adjust the vertical position of the replaced text, ensuring proper alignment with the rest of the document. See Adjust Text Alignment for more details. |
# Search and Replace With Image
Source: https://developer.pdf.co/integrations/zapier/search-and-replace-with-image
Enhance your Zapier workflow by integrating this step to search for text inside a PDF and replace it with an image. This feature supports the use of regular expressions for advanced text search capabilities.
## Input
| Name | Description | Required |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text to Locate** | Specify the text you want to find within the PDF. | Yes |
| **Replacement Image URL** | Provide the URL of the image to be used as a replacement for the located text. | Yes |
| **Enable Regular Expressions** | Activate this option to use regular expressions in your search, such as `[0-9]{3}-[0-9]{2}-[0-9]{4}` for finding SSN patterns. | No |
| **Case-Sensitive Search** | Toggle this for a case-sensitive search. | No |
| **Target Pages** | Indicate the pages to be searched. Enter a comma-separated list of page numbers or ranges. Note: The first page is numbered `0`. For example: `0,1-5,7-`. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ------------------------------------------------------------------------------ |
| `url` | The temporary **URL** on the **PDF.co** file server. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `name` | The name of the file. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `AutoCropImages` | boolean | - | Controls whether to crop empty space around an inserted image. See Crop Empty Space Around Images for more information. |
# Search Text
Source: https://developer.pdf.co/integrations/zapier/search-in-pdf
Integrate this step into your Zapier workflow to search for specific text within PDF documents or scanned images, including support for regular expressions.
## Input
| Name | Description | Required |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **PDF Source Link** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Search Query** | Specify the text you want to search for within the PDF document. | Yes |
| **Enable Regular Expressions** | Activate this option to use regular expressions for more complex search patterns. Example: `[0-9]{3}-[0-9]{2}-[0-9]{4}` to locate an SSN. | No |
| **Page Range** | Indicate the page range for the search. Use a comma-separated list of page numbers or ranges. Note: The first page is numbered `0`. For example: `0,1-5,7-`. | No |
| **Simplify Output** | Enable this for a consolidated output in a single list, which facilitates easier data reuse. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| `url` | The temporary URL on the PDF.co file server. |
| `body` | An object array containing the search results. This is visible only if the `Simplified Object` property is set to `False`. |
| `match1` | An object that includes a single search result. Visible when the `Simplified Object` property is set to `True`. |
| `match2` | An object holding another search result, following the same pattern as `match1` for each output file. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| -------------------------------------------------- | ------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ColumnDetectionMode` | string | Content Groups And Borders | Controls column detection/alignment in PDF table extraction. Refer to [Column Detection Mode](#column-detection-mode) for more information. |
| `DetectionMinNumberOfRows` | integer | `1` | Minimum number of rows to detect in a table |
| `DetectionMinNumberOfColumns` | integer | `1` | Minimum number of columns to detect in a table |
| `DetectionMaxNumberOfInvalidSubsequentRowsAllowed` | integer | `0` | Maximum number of invalid subsequent rows allowed in a table |
| `DetectionMinNumberOfLineBreaksBetweenTables` | integer | `0` | Minimum number of line breaks between tables |
| `EnhanceTableBorders` | boolean | `true` | Enhance table borders or not |
| `OCRDetectPageRotation` | boolean | `false` | Controls whether to detect page rotation in the PDF document when OCR applied. Set to true to detect page rotation. See Support page rotation for more information. |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
| `requestParametersDocument` | string | - | |
| `responseParameters` | object | | |
### Column Detection Mode
This might be case when a document contains a number of overlapping invisible text and vector objects that affect column detection. In this case you may need to fix the wrongly positioned data.
Set the options for your column detection via the following `profiles` parameters:
`ColumnDetectionMode` - available values:
* `ContentGroupsAndBorders` (default, no need to specify)
* `ContentGroups`
* `Borders`
* `BorderedTables`
* `ContentGroupsAI`
```json theme={null}
{
"profiles": "{ 'ColumnDetectionMode': 'ContentGroups' }"
}
```
# Split PDF Into Multiple Files
Source: https://developer.pdf.co/integrations/zapier/split-pdf
Enhance your Zapier workflow by integrating this step to split PDF files into multiple files, based on specific page indexes or ranges.
## Input
| Name | Description | Required |
| ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Page Numbers/Ranges** | Specify a comma-separated list of page number or ranges for processing. Note: The first page is indexed as `1`. Use a dash `-` to specify ranges, e.g., `1,2-5,7-`. To indicate a range from a specific page to the end, use a format like `2-`. Use an asterisk `*` to split all pages into files. | Yes |
| **Base Filename for New PDFs** | Name for the output files. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `url1` | This represents the temporary URL of the output file hosted on the PDF.co file server. |
| `url2` | Similarly, this is the temporary URL for another output file on the PDF.co file server. This pattern is used for all output files. |
| `urls` | This is an array of temporary URLs, each pointing to an output file. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as `base64` format, set this to `base64` |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Split PDF Based on Barcode Search
Source: https://developer.pdf.co/integrations/zapier/split-pdf-by-barcode
Enhance your Zapier workflow by integrating this step to segment PDF documents based on barcode search, creating new, separate PDF files for each identified segment.
## Input
| Name | Description | Required |
| ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Barcode Search String** | Specify the barcode search criteria in the format: `[[barcode:]]`. Example: `[[barcode:qrcode]]`. Refer to [supported barcode types](/api/barcode-reader) for further details. | Yes |
| **Exclude Pages with Identified Barcodes** | Set to `True` to exclude pages containing the identified barcodes. Default is `False`. | No |
| **Base Filename for New PDFs** | Define the base filename for the new PDF files. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `url1` | This represents the temporary URL of the output file hosted on the PDF.co file server. |
| `url2` | Similarly, this is the temporary URL for another output file on the PDF.co file server. This pattern is used for all output files. |
| `urls` | This is an array of temporary URLs, each pointing to an output file. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Split PDF Based on Text Search
Source: https://developer.pdf.co/integrations/zapier/split-pdf-by-text
Enhance your Zapier workflow by integrating this step to segment PDF documents based on text search, including OCR capabilities. This feature is particularly useful for creating new PDF files from sections of the original document, identified through specific text or patterns using regular expressions.
## Input
| Name | Description | Required |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| **Source PDF URL** | Provide the URL to the source PDF document, or a `filetoken://` link from [PDF.co Built-In Files Storage](https://app.pdf.co/files). If you use another cloud service such as **Google Drive** or **Dropbox** ensure the link is publicly accessible. | Yes |
| **Text Search String** | Specify the text string for searching within the PDF pages. | Yes |
| **Enable Case-Sensitive Search** | Activate this to `True` for case-sensitive search. Default is `False`. | No |
| **Enable Regular Expression Search** | Set this to `True` to incorporate regular expressions in your search. The default is `False`. | No |
| **Exclude Pages with Identified Text** | Opt this to `True` to exclude pages where the text is found. Default is `False`. | No |
| **OCR Language** | Select the [OCR language](/api/pdf-make-text-searchable-or-unsearchable) for text recognition in scanned PDFs. Default is English. | No |
| **Base Filename for New PDFs** | Define the base filename for the newly created segmented PDF files. | No |
| **Custom Profiles** | A `JSON` string which adds options for the conversion process. See [Custom Profiles](#custom-profiles) for more. | No |
### Source PDF URL & Google
When using **Google Drive**, it’s typically recommended to choose the **File** option. For more advanced file integration techniques, see [Integrating File Sources with pdf.co](/integrations/zapier/input-file-sources).
## Output
| Name | Description |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `url1` | This represents the temporary URL of the output file hosted on the PDF.co file server. |
| `url2` | Similarly, this is the temporary URL for another output file on the PDF.co file server. This pattern is used for all output files. |
| `urls` | This is an array of temporary URLs, each pointing to an output file. |
| `outputLinkValidTill` | A timestamp which indicates how long the `url` will be available for. |
| `error` | Details of any errors (if any). |
| `status` | The [response status](/api/introduction) code. If all good this will be `200`. |
| `jobId` | The unique identifier for the job. |
| `credits` | The credits spent on the process. |
| `remainingCredits` | The credits left on your account. |
| `duration` | The time it took for the process. |
## Custom profiles
Use Custom [Profiles](/api/profiles) to enhance your workflow with additional processing options. Enter `JSON` configuration to customize OCR settings, output format, text extraction methods, and more.
### Sample JSON
```json theme={null}
{ "ImageOptimizationFormat": "JPEG", "JPEGQuality": 25, "ResampleImages": true, "ResamplingResolution": 120, "GrayscaleImages": true }
```
You can use any regular API parameter from the [API Reference](/api) within Zapier using the `std_params` feature in profiles. The `std_params` enables the definition of regular API parameters in a JSON format, See [Standard Parameters](/api/profiles#standard-parameters) for detailed documentation and examples.
| Parameter | Type | Default | Description |
| ------------------------- | ------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `outputDataFormat` | string | - | If you require your output as base64 format, set this to base64 |
| `DataEncryptionAlgorithm` | string | - | Controls the encryption algorithm used for data encryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataEncryptionKey` | string | - | Controls the encryption key used for data encryption. See User-Controlled Encryption for more information. |
| `DataEncryptionIV` | string | - | Controls the encryption IV used for data encryption. See User-Controlled Encryption for more information. |
| `DataDecryptionAlgorithm` | string | - | Controls the decryption algorithm used for data decryption. See User-Controlled Encryption for more information. The available algorithms are: AES128, AES192, AES256. |
| `DataDecryptionKey` | string | - | Controls the decryption key used for data decryption. See User-Controlled Encryption for more information. |
| `DataDecryptionIV` | string | - | Controls the decryption IV used for data decryption. See User-Controlled Encryption for more information. |
# Convert
Source: https://developer.pdf.co/knowledgebase/convert
This section covers the conversion of PDF files to various formats, including text, images, and more.
# [PDF from HTML](/api/pdf-from-html/convert)
## I want to Prevent Images from Breaking Across Pages in HTML to PDF
You want to generate a PDF from HTML without having images split awkwardly across pages.
### Why This Happens
When converting HTML to PDF, especially in long documents, browsers or PDF renderers may automatically break content (like images or `
`s) across pages. This results in visual issues where a single image may appear partially on one page and continue on the next.
### How to Keep Images (or Elements) Intact
To prevent this, you'll use a CSS style that instructs the PDF renderer not to break certain elements across pages.
### Step-by-Step Fix
Use the `break-inside:avoid` CSS Rule: Add a `"}
```
This will inject the appropriate CSS when converting from HTML.
### For Developers (API)
When using the [/pdf/convert/from/html](/api/pdf-from-html/convert) or [/pdf/convert/from/url](/api/pdf-from-url) endpoints, include the `profiles` parameter like this:
```json theme={null}
{
"profiles": "{'HTMLCodeHeadInject':''}"
}
```
This ensures the CSS rule is applied during PDF generation.
### Applies To:
* [/pdf/convert/from/html](/api/pdf-from-html/convert)
* [/pdf/convert/from/url](/api/pdf-from-url)
### Helpful Tips
* Always test the final PDF to confirm formatting appears as expected.
* Combine this method with margin settings or custom page breaks for better layout control.
## How to Convert HTML to PDF with a Custom Page Size
You want to generate a PDF from HTML with a specific page size—such as **4 inches by 6 inches**—instead of the default A4.
### How to Set a Custom PDF Page Size
To generate a PDF at a custom size, use the `paperSize` parameter in your API call or integration.
Example: to create a **4x6 inch** PDF, set:
```json theme={null}
"paperSize": "4in 6in"
```
You can also use other units:
* `px` → `200px 300px`
* `mm` → `100mm 150mm`
* `cm` → `10cm 15cm`
* `in` → `4in 6in`
### API Example
Here’s a sample API request to convert HTML to a **4x6 inch** PDF:
```
POST /v1/pdf/convert/from/html HTTP/1.1
Host: api.pdf.co
x-api-key: YOUR_API_KEY
Content-Type: application/json
{
"html": "
Hello World!
Go to PDF.co",
"name": "newDocument.pdf",
"mediaType": "print",
"margins": "0",
"paperSize": "4in 6in",
"orientation": "Portrait",
"printBackground": false
}
```
### Zapier
In the HTML to PDF Converter action:
1. Find **Page Size Override**
2. Enter `4in 6in`
### Make / Integromat
In the PDF.co Convert HTML to PDF module:
1. Set **Paper Size** → **Custom**
2. Enter `4in 6in` in the size field
### Helpful Tips
* Always include the unit (e.g., `in`, `mm`, `px`) in `paperSize`
* Use `margins`: `0` for full-page layouts
* Use `printBackground`: `true` if your HTML needs background styles
## How to Keep Form Controls When Converting HTML to PDF
When you convert an HTML page to PDF, **form controls like input fields, checkboxes, and dropdowns appear as static visuals but are not interactive.**
### Why This Happens
The PDF format doesn’t automatically convert HTML form elements into interactive form fields. They’re rendered as plain graphics rather than fillable fields.
### How to Make PDF Forms Fillable
To add interactive form controls, you need to **create form fields in the PDF after converting from HTML.**
**2-Step Workflow:**
1. Convert the HTML to PDF
Use the [/pdf/convert/from/html](/api/pdf-from-html/convert) API or your preferred HTML-to-PDF tool to generate the static PDF.
2. Add Form Fields to the PDF
Use the [/pdf/edit/add](/api/pdf-add) API endpoint to insert form fields like textboxes, checkboxes, and dropdowns into the PDF.
### Platform-Specific Guides
* Zapier: Follow this step-by-step guide → [Create Fillable Forms with PDF.co in Zapier](https://pdf.co/create-fillable-form-using-pdf-co-and-zapier)
* Make/Integromat: Follow this tutorial → [Create Fillable Forms with PDF.co in Make](https://pdf.co/create-fillable-forms-integromat)
### Helpful Tips
* Use PDF.co's built-in [PDF Inspector Tool](https://app.pdf.co/pdf-edit-add-helper) to visually test field placement.
* Use exact coordinates for better precision when placing fields.
## I Want to Create Navigation Links Inside a PDF When Using HTML to PDF
### The Issue
You want to add clickable navigation links inside a PDF—such as a table of contents that jumps to a section, or internal links between parts of a long document—but the links aren’t working as expected.
### Why This Happens
When converting HTML to PDF, some users may not realize that internal anchor links (`#marker`) require properly placed markers (``) in the HTML for the PDF converter to recognize them. Without these, the links won't work in the output file.
### How to Create Navigation Links
To add internal navigation links to your PDF:
1. Add a Marker Where You Want to Jump
Use this in your HTML:
```html theme={null}
```
2. Add a Link to That Marker
Use this in your HTML:
```html theme={null}
Jump to Marker 1
```
### Example: Create a Table of Contents
Here’s a sample HTML that adds navigation to two sections:
```html theme={null}
Navigation:
Section 1 |
Section 2
Section 1
Content for section 1...
Section 2
Content for section 2...
```
### What Happens in the PDF
Once converted using the HTML to PDF tool, the output PDF will preserve these internal links. Clicking on “Section 1” or “Section 2” in the navigation will jump directly to the respective areas—just like in the original HTML document.
### Helpful Tips
Double-check that both the link (`href="#..."`) and the marker (`name="..."`) match exactly.
* Ensure correct HTML syntax—mismatched quotes or missing closing tags can break links.
* Test in a browser first to confirm links are functioning before converting.
## I Want to Avoid Page Breaks in My PDF
You may be experiencing unwanted page breaks in your generated PDF—such as a section of text or a table being split across two pages.
### Why This Happens
By default, PDFs are set to render in A4. When your content exceeds the default page height, it causes automatic page breaks.
### How to Avoid Page Breaks
To prevent this, you can customize the page size using the paperSize parameter. This ensures the page is long enough to hold all your content without forcing a break.
### Example: Use a Custom Page Size
To avoid page breaks and allow more vertical space, increase the page length.
Set the `paperSize` like this:
```json theme={null}
{
"paperSize": "6in 15in"
}
```
This sets the page width to 6 inches and the height to 15 inches—enough space for taller content like long tables or paragraphs.
### Helpful Tips
* Adjust `paperSize` according to your content length.
* Use longer page heights to accommodate full content without breaks.
* If using `async: true`, monitor job status for large files.
## I Want to Force a New Page in the Output PDF
You want to make sure that certain content in your PDF always starts on a **new page**, but currently, everything appears continuously without page breaks.
### Why It Happens
When you're generating PDFs from HTML or HTML templates, content flows naturally like a webpage—**without automatic page breaks**—unless you explicitly tell it where to break using specific CSS rules.
### How to Force a New Page
To force a page break in your PDF output, use the following **CSS property** in your HTML:
```css theme={null}
page-break-after: always;
```
This ensures that any element it's applied to will **force the next content to start on a new page**.
### Examples
#### Example 1 — Using a `
` tag:
```html theme={null}
```
#### Example 2 — Using a `
` tag:
```html theme={null}
```
These tags act as **"invisible" page breaks** in your layout, ensuring that whatever comes next starts on a new page in the final PDF.
### Helpful Tips
* Apply the style to a `